home.social

#llama — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #llama, aggregated by home.social.

fetched live
  1. Minor architectural choices in dense transformers can cut long context performance by up to 47% when combined. New ablation study shows normalization, GQA, pretraining length, and sliding window attention drive most variation across Llama, Qwen, and Olmo families.

    Source: arXiv cs.CL
    arxiv.org/abs/2608.10296

    #MachineLearning #Llama #Qwen

  2. A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.

    * llama.cpp: `Generation: 12.5 t/s`
    * ollama: `eval rate: 10.17 tokens/s`

    This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂

    #AI #LLM #ollama #llama.cpp #localLLM

  3. A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.

    * llama.cpp: `Generation: 12.5 t/s`
    * ollama: `eval rate: 10.17 tokens/s`

    This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂

    #AI #LLM #ollama #llama.cpp #localLLM

  4. A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.

    * llama.cpp: `Generation: 12.5 t/s`
    * ollama: `eval rate: 10.17 tokens/s`

    This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂

    #AI #LLM #ollama #llama.cpp #localLLM

  5. A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.

    * llama.cpp: `Generation: 12.5 t/s`
    * ollama: `eval rate: 10.17 tokens/s`

    This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂

    #AI #LLM #ollama #llama.cpp #localLLM

  6. A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.

    * llama.cpp: `Generation: 12.5 t/s`
    * ollama: `eval rate: 10.17 tokens/s`

    This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂

    #AI #LLM #ollama #llama.cpp #localLLM

  7. Meta Muse Glimmer’s license is Open Source

    Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.

    Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.

    Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?

  8. Meta Muse Glimmer’s license is Open Source

    Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.

    Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.

    Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?

  9. Meta Muse Glimmer’s license is Open Source

    Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.

    Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.

    Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?

  10. Meta Muse Glimmer’s license is Open Source

    Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.

    Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.

    Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?

  11. Meta Muse Glimmer’s license is Open Source

    Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.

    Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.

    Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?

  12. SPECTRA introduces training-free KV cache compression, enabling 4x to 12x compression with near-lossless quality for long-context models like Llama-3.1-8B and Qwen2.5-7B

    Source: arXiv cs.LG
    arxiv.org/abs/2608.07915

    #MachineLearning #Llama

  13. Frontier language models show divergent response modes under steering pressure, with GPT-5 deflecting reasoning disclosure and Claude Opus 4.7 resisting suppression instructions. A linear probe traces the largest behavioral split to Llama’s internals at 0.87 accuracy.

    Source: arXiv cs.AI
    arxiv.org/abs/2608.06578

    #MachineLearning #Claude #Llama

  14. Meta macht KI-Agenten für zuhause fit. Muse Glimmer plant Aufgaben, schreibt Code und korrigiert Fehler selbst – direkt auf deinem PC, ohne Cloud. Das offene 30B-Modell passt dank Quantisierung auf GPUs ab 24 GB. Ein starkes Stück offene Software. #MuseGlimmer #MetaAI #OpenWeights #Llama #AIGeneratedImage

    all-ai.de/news/news26top/meta-

  15. Meta macht KI-Agenten für zuhause fit. Muse Glimmer plant Aufgaben, schreibt Code und korrigiert Fehler selbst – direkt auf deinem PC, ohne Cloud. Das offene 30B-Modell passt dank Quantisierung auf GPUs ab 24 GB. Ein starkes Stück offene Software. #MuseGlimmer #MetaAI #OpenWeights #Llama #AIGeneratedImage

    all-ai.de/news/news26top/meta-

  16. Meta macht KI-Agenten für zuhause fit. Muse Glimmer plant Aufgaben, schreibt Code und korrigiert Fehler selbst – direkt auf deinem PC, ohne Cloud. Das offene 30B-Modell passt dank Quantisierung auf GPUs ab 24 GB. Ein starkes Stück offene Software. #MuseGlimmer #MetaAI #OpenWeights #Llama #AIGeneratedImage

    all-ai.de/news/news26top/meta-

  17. A new pruning method called Whisper preserves output differences to improve LLM sparsification, outperforming Wanda and SparseGPT on Llama 2 and 3.1 models from 7B to 405B parameters

    Source: arXiv cs.LG
    arxiv.org/abs/2608.06630

    #MachineLearning #Llama