home.social

#llmperformance — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #llmperformance, aggregated by home.social.

fetched live
  1. RT @__tinygrad__: Die Benchmark-Ergebnisse liegen vor: GLM-5.2 auf einer $160k Tinybox Pro v2 Black erreicht 119 Tokens pro Sekunde im Single-Use-Modus und 917 im Aggregat-Modus! Die Inbetriebnahme wurde von (einem anderen) GLM-5.2 in nur einer Stunde durchgeführt, sodass definitiv noch weiteres Leistungspotenzial besteht.

    mehr auf Arint.info

    #AIHardware #Benchmark #GLM52 #LLMPerformance #TinyboxPro #TokenSpeed #arint_info

    https://x.com/__tinygrad__/status/2082929188668153915#m

  2. The Illusion of Performance: Why Throughput Obscures LLM Failure

    Are LLM throughput numbers misleading? Learn why goodput is the new standard for measuring real AI performance and user value as of May 2026.

    #llmperformance, #aitechnology, #goodput, #techmetrics, #aiserving

    newsletter.tf/llm-goodput-vs-t

  3. Engineers are moving away from throughput, which counts all data, to goodput, which only counts useful data. This shift helps fix slow AI responses that users cannot actually use.

    #llmperformance, #aitechnology, #goodput, #techmetrics, #aiserving
    newsletter.tf/llm-goodput-vs-t

  4. New research suggests ditching the dream of a single universal AI assistant. By using a Multi‑Connector Protocol, we can orchestrate specialized AI agents and bots that stay in isolated workflows, manage context locally, and boost LLM performance. Discover why modular tool orchestration may be the future of open‑source AI. #MultiConnectorProtocol #SpecializedBots #ToolOrchestration #LLMPerformance

    🔗 aidailypost.com/news/mcp-appro

  5. New research suggests ditching the dream of a single universal AI assistant. By using a Multi‑Connector Protocol, we can orchestrate specialized AI agents and bots that stay in isolated workflows, manage context locally, and boost LLM performance. Discover why modular tool orchestration may be the future of open‑source AI. #MultiConnectorProtocol #SpecializedBots #ToolOrchestration #LLMPerformance

    🔗 aidailypost.com/news/mcp-appro

  6. LM Cache boosts LLM efficiency, scalability, and cost savings by letting the system remember previous outputs and complementing other optimizations. hackernoon.com/optimizing-llm- #llmperformance

  7. LM Cache boosts LLM efficiency, scalability, and cost savings by letting the system remember previous outputs and complementing other optimizations. hackernoon.com/optimizing-llm- #llmperformance