home.social

#tokens — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #tokens, aggregated by home.social.

  1. The all-you-can-eat AI buffet has closed

    This is a phrase which Andrew Tindall uses in this Drum piece to describe the pricing shift already underway in language models. Most consumers are still insulated from a change which is currently directed at enterprise customers and ‘power users’ of Claude:

    Microsoft’s decision to wind down Claude Code licens in parts of the business is the canary in the blood-diamond mine. A company with enough infrastructure and cloud power to make rain nervous is pushing engineers away from a popular coding tool and toward its own stuff. The official line is convergence, but you cannot ignore that the news lands on Microsoft’s fiscal year-end. Meanwhile, GitHub is moving Copilot itself to usage-based billing. Again, presumably after getting jealous of how much cash Claude has been raking in.

    What all AI leaders tried to avoid is the cold reality of the models we have fired people for and plugged into every work process: they are real, useful and expensive AF. It is a much trickier reflection on AI than “AI is fake and useless.” Fake and useless get adopted slowly. Useful and expensive get governed and rationed. Terrible for profits. Meanwhile, every head of tech has independently selected their favorite AI provider, usually with zero business case or P&L analysis.

    It’s something that Justin Pickard has just published about here, based on a contribution to our workshop during the summer. I love this description in particular of the development work that might be prompted by this sudden shift into token scarcity:

    By 2027, a second layer had grown around model use. Balance widgets sat
    in browser windows; cache timers counted down beside reset timers. ‘Ask Later’
    buttons queued prompts, like pet feeders for work the user did not trust themselves to start at the right time. Spreadsheets recorded token burn, task type, cheaper windows. All of it read from figures the platform had not thought to round. The chat window stayed where it was. The counters multiplied around it.4

    #AI #inferenceRationing #politicalEconomy #rationining #tokenEconomics #tokens
  2. Will economic constraints on token use in organisations drive the emergence of norms?

    I’ve been following the token maxxing discourse with interest. Essentially we’ve seen a tendency to equate quantity of tokens used with the extent of AI integration. It’s hard to measure integration so organisations have turned to the proxy of tokens, assuming that the more tokens you are using then the more you are integrating LLMs into your work. The problem with this is two fold:

    • At present tokens are essentially being subsidised by investors in AI labs interested in maximising adoption of the products. The costs of token to the lab are either minimised for the end user or entirely removed from the equation with unmetered access.
    • The assumption that more use = better is obviously untenable with even a rudimentary knowledge of the ethical and epistemological risks of language models. Further, more use of LLMs might be worse for the organisation because it hinders other forms of work which are essential to the organisation’s mission.

    There is a significant shift underway which is going to change how LLMs are used within organisation, summarised here by 404 media:

    The news highlights a major shift in the tech industry and other companies that use AI: the wave of uninhibited AI growth is over. Some AI providers like GitHub are now charging customers per token rather than a flat subscription fee, leading some companies to burn through their tokens. Uber recently capped employees’ use of AI tools like Claude Code and Cursor; that came after Uber told employees to use AI as much as possible and Uber’s CTO said the company had blown its entire AI budget in four months. And Accenture itself reportedly started requiring senior staff to start using AI or risk losing out on promotions.

    I was wrong to believe that model development was flatlining. A week with Claude Fable, the continual development of Claude Opus and my begrudging appreciation of GPT 5.5 leave me persuaded we’ve come along way since GPT 5. However even if the models are getting more capable, to what extent are those capabilities becoming more expensive? I managed to burn through £100+ in five days playing with Claude Fable and I constantly have Opus switched to max now, even when I vaguely know it’s wasteful. There’s a whole style of use which has taken hold here which isn’t sustainable and is increasingly hitting a brick wall.

    For individuals it raises the question of what you’re willing to pay for. I switch to Max plans when I have a special reason to do so but I never keep the subscription any more. I hit the rate limits with Claude so frequently that it’s left me thinking more carefully about what I do want to use models for and what I don’t want to use models for. The same process is inevitably going to take place in organisations I think in the sense of resource constraints necessitating evaluative criteria for desirable and undesirable use of the model.

    In the meantime though I think it’s imperative that we stop universities from sliding into token maxxing with the use of enterprise systems because the entire price model for this is likely to change dramatically in the coming months. Given the wider economics of the industry, will any AI lab really retain per seat pricing for enterprise packages (i.e. paying by user rather than for tokens?) in the longer term? If not then the norms about use we establish now will have significant financial consequences further down the line.

    #AIIntegration #compute #economics #organisations #tokenMaxxing #tokens
  3. Will economic constraints on token use in organisations drive the emergence of norms?

    I’ve been following the token maxxing discourse with interest. Essentially we’ve seen a tendency to equate quantity of tokens used with the extent of AI integration. It’s hard to measure integration so organisations have turned to the proxy of tokens, assuming that the more tokens you are using then the more you are integrating LLMs into your work. The problem with this is two fold:

    • At present tokens are essentially being subsidised by investors in AI labs interested in maximising adoption of the products. The costs of token to the lab are either minimised for the end user or entirely removed from the equation with unmetered access.
    • The assumption that more use = better is obviously untenable with even a rudimentary knowledge of the ethical and epistemological risks of language models. Further, more use of LLMs might be worse for the organisation because it hinders other forms of work which are essential to the organisation’s mission.

    There is a significant shift underway which is going to change how LLMs are used within organisation, summarised here by 404 media:

    The news highlights a major shift in the tech industry and other companies that use AI: the wave of uninhibited AI growth is over. Some AI providers like GitHub are now charging customers per token rather than a flat subscription fee, leading some companies to burn through their tokens. Uber recently capped employees’ use of AI tools like Claude Code and Cursor; that came after Uber told employees to use AI as much as possible and Uber’s CTO said the company had blown its entire AI budget in four months. And Accenture itself reportedly started requiring senior staff to start using AI or risk losing out on promotions.

    I was wrong to believe that model development was flatlining. A week with Claude Fable, the continual development of Claude Opus and my begrudging appreciation of GPT 5.5 leave me persuaded we’ve come along way since GPT 5. However even if the models are getting more capable, to what extent are those capabilities becoming more expensive? I managed to burn through £100+ in five days playing with Claude Fable and I constantly have Opus switched to max now, even when I vaguely know it’s wasteful. There’s a whole style of use which has taken hold here which isn’t sustainable and is increasingly hitting a brick wall.

    For individuals it raises the question of what you’re willing to pay for. I switch to Max plans when I have a special reason to do so but I never keep the subscription any more. I hit the rate limits with Claude so frequently that it’s left me thinking more carefully about what I do want to use models for and what I don’t want to use models for. The same process is inevitably going to take place in organisations I think in the sense of resource constraints necessitating evaluative criteria for desirable and undesirable use of the model. .

    #AIIntegration #compute #economics #organisations #tokenMaxxing #tokens
  4. Will economic constraints on token use in organisations drive the emergence of norms?

    I’ve been following the token maxxing discourse with interest. Essentially we’ve seen a tendency to equate quantity of tokens used with the extent of AI integration. It’s hard to measure integration so organisations have turned to the proxy of tokens, assuming that the more tokens you are using then the more you are integrating LLMs into your work. The problem with this is two fold:

    • At present tokens are essentially being subsidised by investors in AI labs interested in maximising adoption of the products. The costs of token to the lab are either minimised for the end user or entirely removed from the equation with unmetered access.
    • The assumption that more use = better is obviously untenable with even a rudimentary knowledge of the ethical and epistemological risks of language models. Further, more use of LLMs might be worse for the organisation because it hinders other forms of work which are essential to the organisation’s mission.

    There is a significant shift underway which is going to change how LLMs are used within organisation, summarised here by 404 media:

    The news highlights a major shift in the tech industry and other companies that use AI: the wave of uninhibited AI growth is over. Some AI providers like GitHub are now charging customers per token rather than a flat subscription fee, leading some companies to burn through their tokens. Uber recently capped employees’ use of AI tools like Claude Code and Cursor; that came after Uber told employees to use AI as much as possible and Uber’s CTO said the company had blown its entire AI budget in four months. And Accenture itself reportedly started requiring senior staff to start using AI or risk losing out on promotions.

    I was wrong to believe that model development was flatlining. A week with Claude Fable, the continual development of Claude Opus and my begrudging appreciation of GPT 5.5 leave me persuaded we’ve come along way since GPT 5. However even if the models are getting more capable, to what extent are those capabilities becoming more expensive? I managed to burn through £100+ in five days playing with Claude Fable and I constantly have Opus switched to max now, even when I vaguely know it’s wasteful. There’s a whole style of use which has taken hold here which isn’t sustainable and is increasingly hitting a brick wall.

    For individuals it raises the question of what you’re willing to pay for. I switch to Max plans when I have a special reason to do so but I never keep the subscription any more. I hit the rate limits with Claude so frequently that it’s left me thinking more carefully about what I do want to use models for and what I don’t want to use models for. The same process is inevitably going to take place in organisations I think in the sense of resource constraints necessitating evaluative criteria for desirable and undesirable use of the model.

    In the meantime though I think it’s imperative that we stop universities from sliding into token maxxing with the use of enterprise systems because the entire price model for this is likely to change dramatically in the coming months. Given the wider economics of the industry, will any AI lab really retain per seat pricing for enterprise packages (i.e. paying by user rather than for tokens?) in the longer term? If not then the norms about use we establish now will have significant financial consequences further down the line.

    #AIIntegration #compute #economics #organisations #tokenMaxxing #tokens
  5. Gemini «olvida» mucho antes de lo que Google promete: el millón de tokens no aplica al chat

    Usuarios de los planes pagos AI Pro y Ultra denuncian que el chatbot comienza a perder el hilo de la conversación después de apenas 25-30 mensajes, muy lejos del límite de un millón de tokens que Google publicita. La diferencia entre la ventana de contexto del modelo y la del chat nunca se comunica con claridad (Fuente AndroidAutorithy).

    Google vende sus planes pagos de Gemini con una promesa concreta: una ventana de contexto de hasta un millón de tokens, equivalente a 1.500 páginas de texto o 30.000 líneas de código. El problema es que esa cifra no describe lo que le pasa al usuario en una conversación real.

    Usuarios en X y Reddit denunciaron que, si bien los servidores de Gemini pueden efectivamente ingerir un archivo estático masivo en el primer prompt, la memoria conversacional activa —el contexto dinámico del chat— parece estar severamente limitada, cayendo a un tope aproximado de 16.000 tokens, equivalente a unos 25 o 30 mensajes promedio. El resultado es que el modelo sufre de «amnesia» dentro de la misma sesión de chat, olvidando por completo instrucciones anteriores, bloques de código o restricciones que el usuario había establecido al inicio de la conversación.

    La distinción técnica que Google no comunica con suficiente claridad es la diferencia entre la ventana de contexto del modelo y la del chat. En palabras del usuario de X @Soso_fun_yt, mientras el backend puede procesar archivos de gran tamaño en forma estática, la memoria dinámica de la conversación está embotellada en un límite mucho menor, lo que provoca el olvido progresivo. Algunos usuarios señalaron que la plataforma AI Studio sí ofrece la ventana de contexto correcta, pero esa no es la herramienta que usa la mayoría de los suscriptores.

    La analogía que propone el artículo lo dice todo: es como si tu proveedor de internet anunciara una línea de 1 Gbps en su sitio web, sin mencionar en ningún lugar destacado que la velocidad de subida es de apenas 50 Mbps. Google sí publica información técnica sobre tokens de entrada y salida en su documentación para desarrolladores, pero esa información no llega al usuario promedio que paga su suscripción esperando lo que se le prometió.

    Android Authority consultó a Google sobre la discrepancia entre la ventana de contexto del modelo y la del chat, y sobre si planea ofrecer información más prominente al respecto. La compañía no respondió al momento de publicación. Mientras tanto, quienes usan Gemini para proyectos largos o conversaciones técnicas extendidas deberían saber que el millón de tokens es, por ahora, más un horizonte teórico que una realidad práctica de uso diario.

    #AIStudio #AIUltra #chatbot #contextwindow #gemini #GeminiAI #GeminiPro #google #GoogleAI #IA #InteligenciaArtificial #PORTADA #Suscripciones #tokens #Transparencia
  6. 😂 Uber's #AI #budget went poof in four months, so now they're #rationing #AI #tokens like it's the apocalypse! 💸 Because clearly, advanced AI needs to be managed like a #digital #lemonade #stand. 🍋🚗
    simonwillison.net/2026/Jun/3/u #Uber #comedy #HackerNews #ngated

  7. Unlocking Value: Fair Price Discovery, the Role of Market Makers - Getting a token from inception to market is no mean feat and more often takes year... - news.bitcoin.com/unlocking-val #marketmakers(mm) #acherontrading #cryptotrading #marketmakers #wesleypryor #crypto #tokens #op-ed #mm