home.social

#localllms — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #localllms, aggregated by home.social.

fetched live
  1. Learn how KV cache spilling hurts LLM inference, how to detect the cliff, and how to estimate it using real VRAM and throughput signals. hackernoon.com/how-to-find-the #localllms

  2. Learn how KV cache spilling hurts LLM inference, how to detect the cliff, and how to estimate it using real VRAM and throughput signals. hackernoon.com/how-to-find-the #localllms

  3. Learn how KV cache spilling hurts LLM inference, how to detect the cliff, and how to estimate it using real VRAM and throughput signals. hackernoon.com/how-to-find-the #localllms

  4. Learn how KV cache spilling hurts LLM inference, how to detect the cliff, and how to estimate it using real VRAM and throughput signals. hackernoon.com/how-to-find-the

  5. Learn how KV cache spilling hurts LLM inference, how to detect the cliff, and how to estimate it using real VRAM and throughput signals. hackernoon.com/how-to-find-the #localllms

  6. GLM-4.7-Flash is a 30B model with 3B active params, so why did it crawl on my mini PC? Not the weights, not the quant. The attention design, and the runtime. hackernoon.com/a-30b-model-cra #localllms

  7. GLM-4.7-Flash is a 30B model with 3B active params, so why did it crawl on my mini PC? Not the weights, not the quant. The attention design, and the runtime. hackernoon.com/a-30b-model-cra #localllms

  8. GLM-4.7-Flash is a 30B model with 3B active params, so why did it crawl on my mini PC? Not the weights, not the quant. The attention design, and the runtime. hackernoon.com/a-30b-model-cra #localllms

  9. GLM-4.7-Flash is a 30B model with 3B active params, so why did it crawl on my mini PC? Not the weights, not the quant. The attention design, and the runtime. hackernoon.com/a-30b-model-cra

  10. GLM-4.7-Flash is a 30B model with 3B active params, so why did it crawl on my mini PC? Not the weights, not the quant. The attention design, and the runtime. hackernoon.com/a-30b-model-cra #localllms

  11. Your local LLM runs fine until it doesn't. A look at KV cache spilling from VRAM into shared memory, and why it happens silently on Windows. hackernoon.com/why-local-llms- #localllms

  12. Your local LLM runs fine until it doesn't. A look at KV cache spilling from VRAM into shared memory, and why it happens silently on Windows. hackernoon.com/why-local-llms- #localllms

  13. Your local LLM runs fine until it doesn't. A look at KV cache spilling from VRAM into shared memory, and why it happens silently on Windows. hackernoon.com/why-local-llms- #localllms

  14. Your local LLM runs fine until it doesn't. A look at KV cache spilling from VRAM into shared memory, and why it happens silently on Windows. hackernoon.com/why-local-llms-

  15. Your local LLM runs fine until it doesn't. A look at KV cache spilling from VRAM into shared memory, and why it happens silently on Windows. hackernoon.com/why-local-llms- #localllms

  16. cleaning out my computer to make space for #localLLMs. the biggest file on here is a 40GB Twitch VOD of me streaming my first time playing Overwatch 2 for my friends in chat. 4K & 60FPS but 1) i was so bad at the game 2) the OBS sound mixing was horrendous -- you can hardly hear how sad i was about Blizzard killing OW1. there is no way i am deleting this lol. the memories! #rip #overwatch #activision_blizzard

  17. cleaning out my computer to make space for #localLLMs. the biggest file on here is a 40GB Twitch VOD of me streaming my first time playing Overwatch 2 for my friends in chat. 4K & 60FPS but 1) i was so bad at the game 2) the OBS sound mixing was horrendous -- you can hardly hear how sad i was about Blizzard killing OW1. there is no way i am deleting this lol. the memories! #rip #overwatch #activision_blizzard

  18. J'ai pris un moment pour jouer avec les données de #LLM de @betagouv

    👉 comparia.beta.gouv.fr/ranking

    Un modèle local capable consomme 40 à 120x moins d'énergie par token qu'un frontier!!!

    Pour mon cas, utiliser les modèles locaux pour le "quotidien" (rédaction, brainsto, recherches) et garder les API frontier seulement pour le dev, c'est ❗une emprunte carbone divisée par 80❗

    + la confidentialité des données :))

    Note : la conso comparIA est un ordre de grandeur

    #data #AI #localLLMs

  19. J'ai pris un moment pour jouer avec les données de #LLM de @betagouv

    👉 comparia.beta.gouv.fr/ranking

    Un modèle local capable consomme 40 à 120x moins d'énergie par token qu'un frontier!!!

    Pour mon cas, utiliser les modèles locaux pour le "quotidien" (rédaction, brainsto, recherches) et garder les API frontier seulement pour le dev, c'est ❗une emprunte carbone divisée par 80❗

    + la confidentialité des données :))

    Note : la conso comparIA est un ordre de grandeur

    #data #AI #localLLMs

  20. J'ai pris un moment pour jouer avec les données de #LLM de @betagouv

    👉 comparia.beta.gouv.fr/ranking

    Un modèle local capable consomme 40 à 120x moins d'énergie par token qu'un frontier!!!

    Pour mon cas, utiliser les modèles locaux pour le "quotidien" (rédaction, brainsto, recherches) et garder les API frontier seulement pour le dev, c'est ❗une emprunte carbone divisée par 80❗

    + la confidentialité des données :))

    Note : la conso comparIA est un ordre de grandeur

    #data #AI #localLLMs

  21. J'ai pris un moment pour jouer avec les données de #LLM de @betagouv

    👉 comparia.beta.gouv.fr/ranking

    Un modèle local capable consomme 40 à 120x moins d'énergie par token qu'un frontier!!!

    Pour mon cas, utiliser les modèles locaux pour le "quotidien" (rédaction, brainsto, recherches) et garder les API frontier seulement pour le dev, c'est ❗une emprunte carbone divisée par 80❗

    + la confidentialité des données :))

    Note : la conso comparIA est un ordre de grandeur

    #data #AI #localLLMs

  22. J'ai pris un moment pour jouer avec les données de #LLM de @betagouv

    👉 comparia.beta.gouv.fr/ranking

    Un modèle local capable consomme 40 à 120x moins d'énergie par token qu'un frontier!!!

    Pour mon cas, utiliser les modèles locaux pour le "quotidien" (rédaction, brainsto, recherches) et garder les API frontier seulement pour le dev, c'est ❗une emprunte carbone divisée par 80❗

    + la confidentialité des données :))

    Note : la conso comparIA est un ordre de grandeur

    #data #AI #localLLMs

  23. NVIDIA's RTX Spark brings CUDA and 128GB unified memory to mainstream Windows PCs this fall. The move reshapes local-AI hardware decisions: the question shifts from 'buy the biggest card you can afford' to 'which constraint fails first—memory, bandwidth, or model quality.' implicator.ai/nvidias-rtx-spar #AI #Hardware #LocalLLMs

  24. NVIDIA's RTX Spark brings CUDA and 128GB unified memory to mainstream Windows PCs this fall. The move reshapes local-AI hardware decisions: the question shifts from 'buy the biggest card you can afford' to 'which constraint fails first—memory, bandwidth, or model quality.' implicator.ai/nvidias-rtx-spar #AI #Hardware #LocalLLMs

  25. NVIDIA's RTX Spark brings CUDA and 128GB unified memory to mainstream Windows PCs this fall. The move reshapes local-AI hardware decisions: the question shifts from 'buy the biggest card you can afford' to 'which constraint fails first—memory, bandwidth, or model quality.' implicator.ai/nvidias-rtx-spar #AI #Hardware #LocalLLMs

  26. NVIDIA's RTX Spark brings CUDA and 128GB unified memory to mainstream Windows PCs this fall. The move reshapes local-AI hardware decisions: the question shifts from 'buy the biggest card you can afford' to 'which constraint fails first—memory, bandwidth, or model quality.' implicator.ai/nvidias-rtx-spar #AI #Hardware #LocalLLMs

  27. NVIDIA's RTX Spark brings CUDA and 128GB unified memory to mainstream Windows PCs this fall. The move reshapes local-AI hardware decisions: the question shifts from 'buy the biggest card you can afford' to 'which constraint fails first—memory, bandwidth, or model quality.' implicator.ai/nvidias-rtx-spar #AI #Hardware #LocalLLMs

  28. Mein #arbeitgeber labert grade in so nem #MicrosoftTeams Call für alle Mitarbeiter was von #digitalesouveranitat und dann soll ich MEHR mit #microsoft #github #copilot machen. Und selbstverständlich wird ALLES #ai. Sogar unsere TLD wechselt von .net auf .ai.
    Wir sollen ganz explizit doch bitte #ki in die tägliche #Arbeit einbinden, der Vertrieb soll sich beim. Aber bloß "unsere" nutzen, wegen den Daten. Muss mich gleich mal informieren, ob #localLLMs erlaubt sind.
    hessen.social/@Moonstone2487/1

  29. Mein #arbeitgeber labert grade in so nem #MicrosoftTeams Call für alle Mitarbeiter was von #digitalesouveranitat und dann soll ich MEHR mit #microsoft #github #copilot machen. Und selbstverständlich wird ALLES #ai. Sogar unsere TLD wechselt von .net auf .ai.
    Wir sollen ganz explizit doch bitte #ki in die tägliche #Arbeit einbinden, der Vertrieb soll sich beim. Aber bloß "unsere" nutzen, wegen den Daten. Muss mich gleich mal informieren, ob #localLLMs erlaubt sind.
    hessen.social/@Moonstone2487/1

  30. Mein #arbeitgeber labert grade in so nem #MicrosoftTeams Call für alle Mitarbeiter was von #digitalesouveranitat und dann soll ich MEHR mit #microsoft #github #copilot machen. Und selbstverständlich wird ALLES #ai. Sogar unsere TLD wechselt von .net auf .ai.
    Wir sollen ganz explizit doch bitte #ki in die tägliche #Arbeit einbinden, der Vertrieb soll sich beim. Aber bloß "unsere" nutzen, wegen den Daten. Muss mich gleich mal informieren, ob #localLLMs erlaubt sind.
    hessen.social/@Moonstone2487/1

  31. Mein #arbeitgeber labert grade in so nem #MicrosoftTeams Call für alle Mitarbeiter was von #digitalesouveranitat und dann soll ich MEHR mit #microsoft #github #copilot machen. Und selbstverständlich wird ALLES #ai. Sogar unsere TLD wechselt von .net auf .ai.
    Wir sollen ganz explizit doch bitte #ki in die tägliche #Arbeit einbinden, der Vertrieb soll sich beim. Aber bloß "unsere" nutzen, wegen den Daten. Muss mich gleich mal informieren, ob #localLLMs erlaubt sind.
    hessen.social/@Moonstone2487/1

  32. Mein #arbeitgeber labert grade in so nem #MicrosoftTeams Call für alle Mitarbeiter was von #digitalesouveranitat und dann soll ich MEHR mit #microsoft #github #copilot machen. Und selbstverständlich wird ALLES #ai. Sogar unsere TLD wechselt von .net auf .ai.
    Wir sollen ganz explizit doch bitte #ki in die tägliche #Arbeit einbinden, der Vertrieb soll sich beim. Aber bloß "unsere" nutzen, wegen den Daten. Muss mich gleich mal informieren, ob #localLLMs erlaubt sind.
    hessen.social/@Moonstone2487/1