home.social

#tokens — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #tokens, aggregated by home.social.

  1. 🎉 WOW, #RTK claims to magically cut #tokens like a chef on steroids, but apparently our trusty benchmark fairy tale machine says: "Nope, nada, zilch!" 🎩🔧 With a disclaimer hidden in the README like a plot twist, maybe those #GitHub #stars are just as real as unicorns! 🦄✨
    quesma.com/blog/does-rtk-make- #Magic #RTK #Benchmarks #Unicorns #HackerNews #ngated

  2. 🎉 WOW, #RTK claims to magically cut #tokens like a chef on steroids, but apparently our trusty benchmark fairy tale machine says: "Nope, nada, zilch!" 🎩🔧 With a disclaimer hidden in the README like a plot twist, maybe those #GitHub #stars are just as real as unicorns! 🦄✨
    quesma.com/blog/does-rtk-make- #Magic #RTK #Benchmarks #Unicorns #HackerNews #ngated

  3. 🎉 WOW, #RTK claims to magically cut #tokens like a chef on steroids, but apparently our trusty benchmark fairy tale machine says: "Nope, nada, zilch!" 🎩🔧 With a disclaimer hidden in the README like a plot twist, maybe those #GitHub #stars are just as real as unicorns! 🦄✨
    quesma.com/blog/does-rtk-make- #Magic #RTK #Benchmarks #Unicorns #HackerNews #ngated

  4. 🎉 WOW, #RTK claims to magically cut #tokens like a chef on steroids, but apparently our trusty benchmark fairy tale machine says: "Nope, nada, zilch!" 🎩🔧 With a disclaimer hidden in the README like a plot twist, maybe those #GitHub #stars are just as real as unicorns! 🦄✨
    quesma.com/blog/does-rtk-make- #Magic #RTK #Benchmarks #Unicorns #HackerNews #ngated

  5. 🎉 WOW, #RTK claims to magically cut #tokens like a chef on steroids, but apparently our trusty benchmark fairy tale machine says: "Nope, nada, zilch!" 🎩🔧 With a disclaimer hidden in the README like a plot twist, maybe those #GitHub #stars are just as real as unicorns! 🦄✨
    quesma.com/blog/does-rtk-make- #Magic #RTK #Benchmarks #Unicorns #HackerNews #ngated

  6. The all-you-can-eat AI buffet has closed

    This is a phrase which Andrew Tindall uses in this Drum piece to describe the pricing shift already underway in language models. Most consumers are still insulated from a change which is currently directed at enterprise customers and ‘power users’ of Claude:

    Microsoft’s decision to wind down Claude Code licens in parts of the business is the canary in the blood-diamond mine. A company with enough infrastructure and cloud power to make rain nervous is pushing engineers away from a popular coding tool and toward its own stuff. The official line is convergence, but you cannot ignore that the news lands on Microsoft’s fiscal year-end. Meanwhile, GitHub is moving Copilot itself to usage-based billing. Again, presumably after getting jealous of how much cash Claude has been raking in.

    What all AI leaders tried to avoid is the cold reality of the models we have fired people for and plugged into every work process: they are real, useful and expensive AF. It is a much trickier reflection on AI than “AI is fake and useless.” Fake and useless get adopted slowly. Useful and expensive get governed and rationed. Terrible for profits. Meanwhile, every head of tech has independently selected their favorite AI provider, usually with zero business case or P&L analysis.

    It’s something that Justin Pickard has just published about here, based on a contribution to our workshop during the summer. I love this description in particular of the development work that might be prompted by this sudden shift into token scarcity:

    By 2027, a second layer had grown around model use. Balance widgets sat
    in browser windows; cache timers counted down beside reset timers. ‘Ask Later’
    buttons queued prompts, like pet feeders for work the user did not trust themselves to start at the right time. Spreadsheets recorded token burn, task type, cheaper windows. All of it read from figures the platform had not thought to round. The chat window stayed where it was. The counters multiplied around it.4

    #AI #inferenceRationing #politicalEconomy #rationining #tokenEconomics #tokens
  7. The all-you-can-eat AI buffet has closed

    This is a phrase which Andrew Tindall uses in this Drum piece to describe the pricing shift already underway in language models. Most consumers are still insulated from a change which is currently directed at enterprise customers and ‘power users’ of Claude:

    Microsoft’s decision to wind down Claude Code licens in parts of the business is the canary in the blood-diamond mine. A company with enough infrastructure and cloud power to make rain nervous is pushing engineers away from a popular coding tool and toward its own stuff. The official line is convergence, but you cannot ignore that the news lands on Microsoft’s fiscal year-end. Meanwhile, GitHub is moving Copilot itself to usage-based billing. Again, presumably after getting jealous of how much cash Claude has been raking in.

    What all AI leaders tried to avoid is the cold reality of the models we have fired people for and plugged into every work process: they are real, useful and expensive AF. It is a much trickier reflection on AI than “AI is fake and useless.” Fake and useless get adopted slowly. Useful and expensive get governed and rationed. Terrible for profits. Meanwhile, every head of tech has independently selected their favorite AI provider, usually with zero business case or P&L analysis.

    It’s something that Justin Pickard has just published about here, based on a contribution to our workshop during the summer. I love this description in particular of the development work that might be prompted by this sudden shift into token scarcity:

    By 2027, a second layer had grown around model use. Balance widgets sat
    in browser windows; cache timers counted down beside reset timers. ‘Ask Later’
    buttons queued prompts, like pet feeders for work the user did not trust themselves to start at the right time. Spreadsheets recorded token burn, task type, cheaper windows. All of it read from figures the platform had not thought to round. The chat window stayed where it was. The counters multiplied around it.4

    #AI #inferenceRationing #politicalEconomy #rationining #tokenEconomics #tokens
  8. The all-you-can-eat AI buffet has closed

    This is a phrase which Andrew Tindall uses in this Drum piece to describe the pricing shift already underway in language models. Most consumers are still insulated from a change which is currently directed at enterprise customers and ‘power users’ of Claude:

    Microsoft’s decision to wind down Claude Code licens in parts of the business is the canary in the blood-diamond mine. A company with enough infrastructure and cloud power to make rain nervous is pushing engineers away from a popular coding tool and toward its own stuff. The official line is convergence, but you cannot ignore that the news lands on Microsoft’s fiscal year-end. Meanwhile, GitHub is moving Copilot itself to usage-based billing. Again, presumably after getting jealous of how much cash Claude has been raking in.

    What all AI leaders tried to avoid is the cold reality of the models we have fired people for and plugged into every work process: they are real, useful and expensive AF. It is a much trickier reflection on AI than “AI is fake and useless.” Fake and useless get adopted slowly. Useful and expensive get governed and rationed. Terrible for profits. Meanwhile, every head of tech has independently selected their favorite AI provider, usually with zero business case or P&L analysis.

    It’s something that Justin Pickard has just published about here, based on a contribution to our workshop during the summer. I love this description in particular of the development work that might be prompted by this sudden shift into token scarcity:

    By 2027, a second layer had grown around model use. Balance widgets sat
    in browser windows; cache timers counted down beside reset timers. ‘Ask Later’
    buttons queued prompts, like pet feeders for work the user did not trust themselves to start at the right time. Spreadsheets recorded token burn, task type, cheaper windows. All of it read from figures the platform had not thought to round. The chat window stayed where it was. The counters multiplied around it.4

    #AI #inferenceRationing #politicalEconomy #rationining #tokenEconomics #tokens
  9. The all-you-can-eat AI buffet has closed

    This is a phrase which Andrew Tindall uses in this Drum piece to describe the pricing shift already underway in language models. Most consumers are still insulated from a change which is currently directed at enterprise customers and ‘power users’ of Claude:

    Microsoft’s decision to wind down Claude Code licens in parts of the business is the canary in the blood-diamond mine. A company with enough infrastructure and cloud power to make rain nervous is pushing engineers away from a popular coding tool and toward its own stuff. The official line is convergence, but you cannot ignore that the news lands on Microsoft’s fiscal year-end. Meanwhile, GitHub is moving Copilot itself to usage-based billing. Again, presumably after getting jealous of how much cash Claude has been raking in.

    What all AI leaders tried to avoid is the cold reality of the models we have fired people for and plugged into every work process: they are real, useful and expensive AF. It is a much trickier reflection on AI than “AI is fake and useless.” Fake and useless get adopted slowly. Useful and expensive get governed and rationed. Terrible for profits. Meanwhile, every head of tech has independently selected their favorite AI provider, usually with zero business case or P&L analysis.

    It’s something that Justin Pickard has just published about here, based on a contribution to our workshop during the summer. I love this description in particular of the development work that might be prompted by this sudden shift into token scarcity:

    By 2027, a second layer had grown around model use. Balance widgets sat
    in browser windows; cache timers counted down beside reset timers. ‘Ask Later’
    buttons queued prompts, like pet feeders for work the user did not trust themselves to start at the right time. Spreadsheets recorded token burn, task type, cheaper windows. All of it read from figures the platform had not thought to round. The chat window stayed where it was. The counters multiplied around it.4

    #AI #inferenceRationing #politicalEconomy #rationining #tokenEconomics #tokens
  10. The all-you-can-eat AI buffet has closed

    This is a phrase which Andrew Tindall uses in this Drum piece to describe the pricing shift already underway in language models. Most consumers are still insulated from a change which is currently directed at enterprise customers and ‘power users’ of Claude:

    Microsoft’s decision to wind down Claude Code licens in parts of the business is the canary in the blood-diamond mine. A company with enough infrastructure and cloud power to make rain nervous is pushing engineers away from a popular coding tool and toward its own stuff. The official line is convergence, but you cannot ignore that the news lands on Microsoft’s fiscal year-end. Meanwhile, GitHub is moving Copilot itself to usage-based billing. Again, presumably after getting jealous of how much cash Claude has been raking in.

    What all AI leaders tried to avoid is the cold reality of the models we have fired people for and plugged into every work process: they are real, useful and expensive AF. It is a much trickier reflection on AI than “AI is fake and useless.” Fake and useless get adopted slowly. Useful and expensive get governed and rationed. Terrible for profits. Meanwhile, every head of tech has independently selected their favorite AI provider, usually with zero business case or P&L analysis.

    It’s something that Justin Pickard has just published about here, based on a contribution to our workshop during the summer. I love this description in particular of the development work that might be prompted by this sudden shift into token scarcity:

    By 2027, a second layer had grown around model use. Balance widgets sat
    in browser windows; cache timers counted down beside reset timers. ‘Ask Later’
    buttons queued prompts, like pet feeders for work the user did not trust themselves to start at the right time. Spreadsheets recorded token burn, task type, cheaper windows. All of it read from figures the platform had not thought to round. The chat window stayed where it was. The counters multiplied around it.4

    #AI #inferenceRationing #politicalEconomy #rationining #tokenEconomics #tokens
  11. Gemini «olvida» mucho antes de lo que Google promete: el millón de tokens no aplica al chat

    Usuarios de los planes pagos AI Pro y Ultra denuncian que el chatbot comienza a perder el hilo de la conversación después de apenas 25-30 mensajes, muy lejos del límite de un millón de tokens que Google publicita. La diferencia entre la ventana de contexto del modelo y la del chat nunca se comunica con claridad (Fuente AndroidAutorithy).

    Google vende sus planes pagos de Gemini con una promesa concreta: una ventana de contexto de hasta un millón de tokens, equivalente a 1.500 páginas de texto o 30.000 líneas de código. El problema es que esa cifra no describe lo que le pasa al usuario en una conversación real.

    Usuarios en X y Reddit denunciaron que, si bien los servidores de Gemini pueden efectivamente ingerir un archivo estático masivo en el primer prompt, la memoria conversacional activa —el contexto dinámico del chat— parece estar severamente limitada, cayendo a un tope aproximado de 16.000 tokens, equivalente a unos 25 o 30 mensajes promedio. El resultado es que el modelo sufre de «amnesia» dentro de la misma sesión de chat, olvidando por completo instrucciones anteriores, bloques de código o restricciones que el usuario había establecido al inicio de la conversación.

    La distinción técnica que Google no comunica con suficiente claridad es la diferencia entre la ventana de contexto del modelo y la del chat. En palabras del usuario de X @Soso_fun_yt, mientras el backend puede procesar archivos de gran tamaño en forma estática, la memoria dinámica de la conversación está embotellada en un límite mucho menor, lo que provoca el olvido progresivo. Algunos usuarios señalaron que la plataforma AI Studio sí ofrece la ventana de contexto correcta, pero esa no es la herramienta que usa la mayoría de los suscriptores.

    La analogía que propone el artículo lo dice todo: es como si tu proveedor de internet anunciara una línea de 1 Gbps en su sitio web, sin mencionar en ningún lugar destacado que la velocidad de subida es de apenas 50 Mbps. Google sí publica información técnica sobre tokens de entrada y salida en su documentación para desarrolladores, pero esa información no llega al usuario promedio que paga su suscripción esperando lo que se le prometió.

    Android Authority consultó a Google sobre la discrepancia entre la ventana de contexto del modelo y la del chat, y sobre si planea ofrecer información más prominente al respecto. La compañía no respondió al momento de publicación. Mientras tanto, quienes usan Gemini para proyectos largos o conversaciones técnicas extendidas deberían saber que el millón de tokens es, por ahora, más un horizonte teórico que una realidad práctica de uso diario.

    #AIStudio #AIUltra #chatbot #contextwindow #gemini #GeminiAI #GeminiPro #google #GoogleAI #IA #InteligenciaArtificial #PORTADA #Suscripciones #tokens #Transparencia