home.social

#lmcache — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #lmcache, aggregated by home.social.

fetched live
  1. RT @JafarNajafov: TRANSLASATION: So beschleunigst du deine LLM um das 3- bis 10-fache! (100% Open-Source) Es heißt LMCache. Eine KV-Cache-Schicht, die wiederverwendbare Texte über GPU, CPU, Festplatte und sogar S3 speichert und sie dann in jeder vLLM- oder SGLang-Instanz wiederverwendet. Nicht nur Prefix-Caching. Jeder wiederverwendbare Text, überall im Prompt, auf jedem Knoten. In Kombination mit vLLM erzielen Teams eine 3- bis 10-fach niedrigere TTFT und massive Einsparungen bei GPU-Zyklen bei Multi-Round-QA- und RAG-Workloads. Bereits von Google Cloud, CoreWeave, GMI Cloud, Redis, Weka und NVIDIA Dynamo übernommen. Unter der Apache 2.0-Lizenz. Installation in einer Zeile: pip install lmcache Die Inferenz wird bald viel günstiger. github.com/LMCache/LMCache

    mehr auf Arint.info

    #AI #GPU #Inference #LLM #LMCache #OpenSource #arint_info

    https://x.com/JafarNajafov/status/2078065111584186865#m

  2. RT @JafarNajafov: TRANSLASATION: So beschleunigst du deine LLM um das 3- bis 10-fache! (100% Open-Source) Es heißt LMCache. Eine KV-Cache-Schicht, die wiederverwendbare Texte über GPU, CPU, Festplatte und sogar S3 speichert und sie dann in jeder vLLM- oder SGLang-Instanz wiederverwendet. Nicht nur Prefix-Caching. Jeder wiederverwendbare Text, überall im Prompt, auf jedem Knoten. In Kombination mit vLLM erzielen Teams eine 3- bis 10-fach niedrigere TTFT und massive Einsparungen bei GPU-Zyklen bei Multi-Round-QA- und RAG-Workloads. Bereits von Google Cloud, CoreWeave, GMI Cloud, Redis, Weka und NVIDIA Dynamo übernommen. Unter der Apache 2.0-Lizenz. Installation in einer Zeile: pip install lmcache Die Inferenz wird bald viel günstiger. github.com/LMCache/LMCache

    mehr auf Arint.info

    #AI #GPU #Inference #LLM #LMCache #OpenSource #arint_info

    https://x.com/JafarNajafov/status/2078065111584186865#m

  3. Do you want to compare the caching performance of your LLM serving stack? We've put together a simple command line tool to do so. Introducing Tensormesh Benchmark.
    tensormesh.ai/blog-posts/tenso

    #llm #ai #kvcache #lmcache #vllm #benchmarking

  4. 🚀 Behold, the magical #LMCache that promises to triple your LLM's #throughput, as if by waving a wand made of #Redis and marketing buzzwords. 🤖✨ But wait, there's more! Experience the thrill of saving milliseconds while drowning in GitHub's relentless onslaught of #features you never asked for. 🤯🙄
    github.com/LMCache/LMCache #LLM #GitHub #Innovation #HackerNews #ngated

  5. 🚀 Behold, the magical #LMCache that promises to triple your LLM's #throughput, as if by waving a wand made of #Redis and marketing buzzwords. 🤖✨ But wait, there's more! Experience the thrill of saving milliseconds while drowning in GitHub's relentless onslaught of #features you never asked for. 🤯🙄
    github.com/LMCache/LMCache #LLM #GitHub #Innovation #HackerNews #ngated