#llama — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #llama, aggregated by home.social.
-
Minor architectural choices in dense transformers can cut long context performance by up to 47% when combined. New ablation study shows normalization, GQA, pretraining length, and sliding window attention drive most variation across Llama, Qwen, and Olmo families.
Source: arXiv cs.CL
https://arxiv.org/abs/2608.10296 -
A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.
* llama.cpp: `Generation: 12.5 t/s`
* ollama: `eval rate: 10.17 tokens/s`This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂
-
A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.
* llama.cpp: `Generation: 12.5 t/s`
* ollama: `eval rate: 10.17 tokens/s`This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂
-
A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.
* llama.cpp: `Generation: 12.5 t/s`
* ollama: `eval rate: 10.17 tokens/s`This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂
-
A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.
* llama.cpp: `Generation: 12.5 t/s`
* ollama: `eval rate: 10.17 tokens/s`This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂
-
A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.
* llama.cpp: `Generation: 12.5 t/s`
* ollama: `eval rate: 10.17 tokens/s`This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂
-
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
-
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
-
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
-
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
-
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
-
Meta Muse Glimmer’s license is Open Source
Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.
Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.
Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?
-
Meta Muse Glimmer’s license is Open Source
Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.
Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.
Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?
-
Meta Muse Glimmer’s license is Open Source
Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.
Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.
Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?
-
Meta Muse Glimmer’s license is Open Source
Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.
Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.
Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?
-
Meta Muse Glimmer’s license is Open Source
Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.
Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.
Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?
-
SPECTRA introduces training-free KV cache compression, enabling 4x to 12x compression with near-lossless quality for long-context models like Llama-3.1-8B and Qwen2.5-7B
Source: arXiv cs.LG
https://arxiv.org/abs/2608.07915 -
Frontier language models show divergent response modes under steering pressure, with GPT-5 deflecting reasoning disclosure and Claude Opus 4.7 resisting suppression instructions. A linear probe traces the largest behavioral split to Llama’s internals at 0.87 accuracy.
Source: arXiv cs.AI
https://arxiv.org/abs/2608.06578 -
Meta macht KI-Agenten für zuhause fit. Muse Glimmer plant Aufgaben, schreibt Code und korrigiert Fehler selbst – direkt auf deinem PC, ohne Cloud. Das offene 30B-Modell passt dank Quantisierung auf GPUs ab 24 GB. Ein starkes Stück offene Software. #MuseGlimmer #MetaAI #OpenWeights #Llama #AIGeneratedImage
https://www.all-ai.de/news/news26top/meta-muse-glimmer-agent
-
Meta macht KI-Agenten für zuhause fit. Muse Glimmer plant Aufgaben, schreibt Code und korrigiert Fehler selbst – direkt auf deinem PC, ohne Cloud. Das offene 30B-Modell passt dank Quantisierung auf GPUs ab 24 GB. Ein starkes Stück offene Software. #MuseGlimmer #MetaAI #OpenWeights #Llama #AIGeneratedImage
https://www.all-ai.de/news/news26top/meta-muse-glimmer-agent
-
Meta macht KI-Agenten für zuhause fit. Muse Glimmer plant Aufgaben, schreibt Code und korrigiert Fehler selbst – direkt auf deinem PC, ohne Cloud. Das offene 30B-Modell passt dank Quantisierung auf GPUs ab 24 GB. Ein starkes Stück offene Software. #MuseGlimmer #MetaAI #OpenWeights #Llama #AIGeneratedImage
https://www.all-ai.de/news/news26top/meta-muse-glimmer-agent
-
A new pruning method called Whisper preserves output differences to improve LLM sparsification, outperforming Wanda and SparseGPT on Llama 2 and 3.1 models from 7B to 405B parameters
Source: arXiv cs.LG
https://arxiv.org/abs/2608.06630 -
Building a Rust Inference Engine That Matches Llama.cpp
https://www.fratepietro.com/2026/ferrox-rust-gguf-inference-engine/
-
Building a Rust Inference Engine That Matches Llama.cpp
https://www.fratepietro.com/2026/ferrox-rust-gguf-inference-engine/
-
Building a Rust Inference Engine That Matches Llama.cpp
https://www.fratepietro.com/2026/ferrox-rust-gguf-inference-engine/
-
Building a Rust Inference Engine That Matches Llama.cpp
https://www.fratepietro.com/2026/ferrox-rust-gguf-inference-engine/
-
Building a Rust Inference Engine That Matches Llama.cpp
https://www.fratepietro.com/2026/ferrox-rust-gguf-inference-engine/
-
CVE Alert: CVE-2026-70640 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70640-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70640 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70640 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70640-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70640 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70640 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70640-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70640 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70640 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70640-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70640 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70640 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70640-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70640 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70638 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70638-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70638 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70638 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70638-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70638 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70638 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70638-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70638 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70638 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70638-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70638 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70638 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70638-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70638 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43632 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43632-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43632 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43632 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43632-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43632 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43632 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43632-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43632 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43632 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43632-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43632 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43632 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43632-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43632 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43629 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43629-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43629 #ggml-org #llama-cpp