home.social

#costoptimization — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #costoptimization, aggregated by home.social.

fetched live
  1. #JetBrains centralised AI usage after development-related spending increased roughly tenfold in 6 months.

    Instead of limiting engineers to approved tools, it built a shared access and accounting layer to preserve tool choice while improving visibility and control over AI consumption.

    Read the full story on #InfoQ 👉 bit.ly/4x6gDO4

    #AI #CostOptimization #FinOps #SoftwareArchitecture

  2. centralised AI usage after development-related spending increased roughly tenfold in 6 months.

    Instead of limiting engineers to approved tools, it built a shared access and accounting layer to preserve tool choice while improving visibility and control over AI consumption.

    Read the full story on 👉 bit.ly/4x6gDO4

  3. AI inference is expensive - but it doesn’t have to be.

    In this #InfoQ talk, Meryem Arik breaks down how to systematically reduce cost per token across AI workloads, with real-world examples from data transformation, offline agents, and aggregated insights.

    🔹 Measure and optimize inference costs across NVIDIA and AMD GPUs
    🔹 Understand key tradeoffs in vLLM, SGLang, and Dynamo

    🔗 Watch now: bit.ly/4xyPRNN

    #AI #LLM #CostOptimization

  4. AI inference is expensive - but it doesn’t have to be.

    In this talk, Meryem Arik breaks down how to systematically reduce cost per token across AI workloads, with real-world examples from data transformation, offline agents, and aggregated insights.

    🔹 Measure and optimize inference costs across NVIDIA and AMD GPUs
    🔹 Understand key tradeoffs in vLLM, SGLang, and Dynamo

    🔗 Watch now: bit.ly/4xyPRNN

  5. Uber's GOGCTunner dynamically tunes Go garbage collection against live cgroup memory limits and object utilization, reclaiming 70,000 CPU cores across 30 core services.

    Discover how Uber's "Zero Growth Stack" strategy helped them optimize infrastructure and cap AI coding costs at $1,500/developer after monthly spend peaked at $2,000.

    Read the full story on #InfoQ 👉 bit.ly/3RCEZ2m

    #DevOps #GoLang #AIEngineering #CostOptimization

  6. Uber's GOGCTunner dynamically tunes Go garbage collection against live cgroup memory limits and object utilization, reclaiming 70,000 CPU cores across 30 core services.

    Discover how Uber's "Zero Growth Stack" strategy helped them optimize infrastructure and cap AI coding costs at $1,500/developer after monthly spend peaked at $2,000.

    Read the full story on 👉 bit.ly/3RCEZ2m

  7. 🔄 NadirRouter/NadirClaw

    Routes prompts to the cheapest viable model, verifies output, and escalates if needed, cutting AI costs by 40-70%

    ⭐ Stars: 619
    📅 Last Update: Jul 23, 2026

    github.com/NadirRouter/NadirCl

    #selfhosted #homelab #selfhost #selfhosting #opensource #llm #costoptimization

  8. RT @Teknium: Ich nehme eine Änderung am Hermes Agent Curator vor, die dazu führt, dass standardmäßig nur ungenutzte Skills bereinigt werden. Die Zusammenführung von Skills erfolgt nicht mehr automatisch, es sei denn, man aktiviert diese Funktion über die Konfiguration oder das Dashboard.

    mehr auf Arint.info

    #AIConfiguration #CostOptimization #HermesAgent #LLMOps #SkillManagement #TechUpdate #arint_info

    https://x.com/Teknium/status/2067222815913447678#m