#costoptimization — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #costoptimization, aggregated by home.social.
-
#JetBrains centralised AI usage after development-related spending increased roughly tenfold in 6 months.
Instead of limiting engineers to approved tools, it built a shared access and accounting layer to preserve tool choice while improving visibility and control over AI consumption.
Read the full story on #InfoQ 👉 https://bit.ly/4x6gDO4
-
#JetBrains centralised AI usage after development-related spending increased roughly tenfold in 6 months.
Instead of limiting engineers to approved tools, it built a shared access and accounting layer to preserve tool choice while improving visibility and control over AI consumption.
Read the full story on #InfoQ 👉 https://bit.ly/4x6gDO4
-
AI inference is expensive - but it doesn’t have to be.
In this #InfoQ talk, Meryem Arik breaks down how to systematically reduce cost per token across AI workloads, with real-world examples from data transformation, offline agents, and aggregated insights.
🔹 Measure and optimize inference costs across NVIDIA and AMD GPUs
🔹 Understand key tradeoffs in vLLM, SGLang, and Dynamo🔗 Watch now: https://bit.ly/4xyPRNN
-
AI inference is expensive - but it doesn’t have to be.
In this #InfoQ talk, Meryem Arik breaks down how to systematically reduce cost per token across AI workloads, with real-world examples from data transformation, offline agents, and aggregated insights.
🔹 Measure and optimize inference costs across NVIDIA and AMD GPUs
🔹 Understand key tradeoffs in vLLM, SGLang, and Dynamo🔗 Watch now: https://bit.ly/4xyPRNN
-
Uber's GOGCTunner dynamically tunes Go garbage collection against live cgroup memory limits and object utilization, reclaiming 70,000 CPU cores across 30 core services.
Discover how Uber's "Zero Growth Stack" strategy helped them optimize infrastructure and cap AI coding costs at $1,500/developer after monthly spend peaked at $2,000.
Read the full story on #InfoQ 👉 https://bit.ly/3RCEZ2m
-
Uber's GOGCTunner dynamically tunes Go garbage collection against live cgroup memory limits and object utilization, reclaiming 70,000 CPU cores across 30 core services.
Discover how Uber's "Zero Growth Stack" strategy helped them optimize infrastructure and cap AI coding costs at $1,500/developer after monthly spend peaked at $2,000.
Read the full story on #InfoQ 👉 https://bit.ly/3RCEZ2m
-
🔄 NadirRouter/NadirClaw
Routes prompts to the cheapest viable model, verifies output, and escalates if needed, cutting AI costs by 40-70%
⭐ Stars: 619
📅 Last Update: Jul 23, 2026https://github.com/NadirRouter/NadirClaw
#selfhosted #homelab #selfhost #selfhosting #opensource #llm #costoptimization
-
RT @DataChaz: BIS ZU 95% REDUZIERUNG DER TOKEN-KOSTEN OHNE CODEÄNDERUNGEN
mehr auf Arint.info
-
RT @Teknium: Ich nehme eine Änderung am Hermes Agent Curator vor, die dazu führt, dass standardmäßig nur ungenutzte Skills bereinigt werden. Die Zusammenführung von Skills erfolgt nicht mehr automatisch, es sei denn, man aktiviert diese Funktion über die Konfiguration oder das Dashboard.
mehr auf Arint.info
#AIConfiguration #CostOptimization #HermesAgent #LLMOps #SkillManagement #TechUpdate #arint_info