#llama — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #llama, aggregated by home.social.
-
Minor architectural choices in dense transformers can cut long context performance by up to 47% when combined. New ablation study shows normalization, GQA, pretraining length, and sliding window attention drive most variation across Llama, Qwen, and Olmo families.
Source: arXiv cs.CL
https://arxiv.org/abs/2608.10296 -
A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.
* llama.cpp: `Generation: 12.5 t/s`
* ollama: `eval rate: 10.17 tokens/s`This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂
-
A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.
* llama.cpp: `Generation: 12.5 t/s`
* ollama: `eval rate: 10.17 tokens/s`This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂
-
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
-
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
https://github.com/trycua/cua/blob/main/blog/gpu-passthrough-macos-vms.md
-
Meta Muse Glimmer’s license is Open Source
Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.
Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.
Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?
-
Meta Muse Glimmer’s license is Open Source
Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.
Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.
Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?
-
SPECTRA introduces training-free KV cache compression, enabling 4x to 12x compression with near-lossless quality for long-context models like Llama-3.1-8B and Qwen2.5-7B
Source: arXiv cs.LG
https://arxiv.org/abs/2608.07915 -
Frontier language models show divergent response modes under steering pressure, with GPT-5 deflecting reasoning disclosure and Claude Opus 4.7 resisting suppression instructions. A linear probe traces the largest behavioral split to Llama’s internals at 0.87 accuracy.
Source: arXiv cs.AI
https://arxiv.org/abs/2608.06578 -
Meta macht KI-Agenten für zuhause fit. Muse Glimmer plant Aufgaben, schreibt Code und korrigiert Fehler selbst – direkt auf deinem PC, ohne Cloud. Das offene 30B-Modell passt dank Quantisierung auf GPUs ab 24 GB. Ein starkes Stück offene Software. #MuseGlimmer #MetaAI #OpenWeights #Llama #AIGeneratedImage
https://www.all-ai.de/news/news26top/meta-muse-glimmer-agent
-
A new pruning method called Whisper preserves output differences to improve LLM sparsification, outperforming Wanda and SparseGPT on Llama 2 and 3.1 models from 7B to 405B parameters
Source: arXiv cs.LG
https://arxiv.org/abs/2608.06630 -
Building a Rust Inference Engine That Matches Llama.cpp
https://www.fratepietro.com/2026/ferrox-rust-gguf-inference-engine/
-
Building a Rust Inference Engine That Matches Llama.cpp
https://www.fratepietro.com/2026/ferrox-rust-gguf-inference-engine/
-
CVE Alert: CVE-2026-70640 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70640-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70640 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70640 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70640-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70640 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70638 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70638-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70638 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-70638 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-70638-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-70638 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43632 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43632-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43632 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43632 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43632-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43632 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43629 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43629-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43629 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43629 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43629-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43629 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43627 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43627-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43627 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43628 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43628-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43628 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43627 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43627-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43627 #ggml-org #llama-cpp
-
CVE Alert: CVE-2026-43628 - ggml-org - llama.cpp - https://www.redpacketsecurity.com/cve-alert-cve-2026-43628-ggml-org-llama-cpp/
#OSINT #ThreatIntel #CyberSecurity #cve-2026-43628 #ggml-org #llama-cpp
-
從奈洛比到雅加達,開發者一面倒向中國開源AI模型——便宜好用,還不會被美國關掉TNL國際編譯 2026-08-07 14:34:00 CST
從非洲到拉美,開發者正改用中國開源AI:免費、可改、能自己下載。6月美國一紙禁令讓Anthropic全球撤下Fable模型,把「便宜」推成主權保險。美中模型爭奪,台灣早已選邊。
https://www.thenewslens.com/article/269549
#開源模型 #烏干達 #Pax Silica #OpenAI #Meta #美中科技戰 #奈洛比 #月之暗面 #肯亞 #WAICO #DeepSeek #AI模型 #主權AI #阿里巴巴 #Gemma #Hugging Face #OpenRouter #封閉模型 #Llama #Google #Anthropic #矽盛世 #科技 #全球南方 -
從奈洛比到雅加達,開發者一面倒向中國開源AI模型——便宜好用,還不會被美國關掉TNL國際編譯 2026-08-07 14:34:00 CST
從非洲到拉美,開發者正改用中國開源AI:免費、可改、能自己下載。6月美國一紙禁令讓Anthropic全球撤下Fable模型,把「便宜」推成主權保險。美中模型爭奪,台灣早已選邊。
https://www.thenewslens.com/article/269549
#開源模型 #烏干達 #Pax Silica #OpenAI #Meta #美中科技戰 #奈洛比 #月之暗面 #肯亞 #WAICO #DeepSeek #AI模型 #主權AI #阿里巴巴 #Gemma #Hugging Face #OpenRouter #封閉模型 #Llama #Google #Anthropic #矽盛世 #科技 #全球南方 -
AMD Acquires Taalas. Startup that demoed Llama 3.1 8B at 17k tok/s
-
AMD Acquires Taalas. Startup that demoed Llama 3.1 8B at 17k tok/s
-
heyo~! i'm an #asexual #japanese 19yo #transwoman from #canada that likes to beg everyone for #moderator and pretend to be different people and im proud to tell you all that i have so many alt accounts here on the fediverse i even make #dav1d shit his undersized and already stained #pants! 😏 i am a master of using many words to say nothing and i love to repost my nothing burger #blog posts every time i abandon my previous sock puppets and say that #work was stolen from me by my previous #account. i love being a #drama #llama on here and thank you all for being so kind to me! i promise to steal your works in the future as well and claim them as mine, because they in fact ARE mine! stop smoking #weed you crackheads! thats all for my #introduction post! thanks for all the #kindness! ☺️
-
heyo~! i'm an #asexual #japanese 19yo #transwoman from #canada that likes to beg everyone for #moderator and pretend to be different people and im proud to tell you all that i have so many alt accounts here on the fediverse i even make #dav1d shit his undersized and already stained #pants! 😏 i am a master of using many words to say nothing and i love to repost my nothing burger #blog posts every time i abandon my previous sock puppets and say that #work was stolen from me by my previous #account. i love being a #drama #llama on here and thank you all for being so kind to me! i promise to steal your works in the future as well and claim them as mine, because they in fact ARE mine! stop smoking #weed you crackheads! thats all for my #introduction post! thanks for all the #kindness! ☺️
-
How #China’s #AI Is Surging Across #Africa
At first, it seemed a losing bet. #UnitedStates dominated #opensource AI Meta offered #Llama that developers used worldwide. Then Meta turned to closed models.
Breakthroughs came quickly. In Dec 2024 #DeepSeek AI model matched best models at fraction of cost. Last month, Chinese #MoonshotAI released a model with coding abilities approaching leading #US system.
In #Kenya developers embraced Chinese models.
https://www.nytimes.com/2026/08/05/technology/ai-china-africa.html
https://archive.ph/vz2GN -
How #China’s #AI Is Surging Across #Africa
At first, it seemed a losing bet. #UnitedStates dominated #opensource AI Meta offered #Llama that developers used worldwide. Then Meta turned to closed models.
Breakthroughs came quickly. In Dec 2024 #DeepSeek AI model matched best models at fraction of cost. Last month, Chinese #MoonshotAI released a model with coding abilities approaching leading #US system.
In #Kenya developers embraced Chinese models.
https://www.nytimes.com/2026/08/05/technology/ai-china-africa.html
https://archive.ph/vz2GN -
DeepSeek V4 Flash is no longer text-only 👀
In our internal benchmarks, it delivered significantly better price-performance than others in its class.
We’ve added vision for screen-level understanding, capabilities needed for WebBrain
Available on HF:
https://huggingface.co/webbrain-one/DeepSeek-V4-Flash-0731-Vision-NVFP4 -
DeepSeek V4 Flash is no longer text-only 👀
In our internal benchmarks, it delivered significantly better price-performance than others in its class.
We’ve added vision for screen-level understanding, capabilities needed for WebBrain
Available on HF:
https://huggingface.co/webbrain-one/DeepSeek-V4-Flash-0731-Vision-NVFP4 -
RT @lukepm: I wanted a practical way for my 5-person team to run DeepSeek V4 Flash 0731 with Hermes Agent on 2x RTX PRO 6000 Blackwell GPUs. So I ran 80 fresh tests across llama.cpp, DSpark, and custom vLLM with and without CPU KV offload. Two GPUs work, but KV is the catch. 👇 THE BENCHMARK I cared about team use, not one flashy speed number. Seven people can hit the server at the same time. Hermes Agent can also create long, tool-heavy conversations. So I tested both speed and concurrency. Hardware: 2x NVIDIA RTX PRO 6000 Blackwell 96GB over PCIe. No NVLink. Prompts were exactly 2K, 32K, 64K, and 100K tokens. Concurrency was 1, 4, 8, 16, and 32. I used a fixed 13-record Spec-Bench subset. Each request generated 128 tokens with greedy streaming. GPU prompt caching was off. For CPU offload, unique early prompt blocks kept external cache reuse at 0%. All speeds are tokens per second. Prefill and effective aggregate decode are shown separately. 1. LLAMA.CPP WITHOUT DSPARK Target: DeepSeek V4 Flash 0731 UD-Q8_K_XL GGUF by Unsloth. KV cache: F16. Slots: 32, with up to 128K context per slot. This used llama.cpp layer split. Complete model layers were divided between the GPUs. The tensor-parallel-like row split mode does not support DeepSeek V4 Flash 0731 yet. The command still uses --tensor-split 1,1. In layer-split mode, that only sets the GPU allocation ratio. It does not enable tensor parallelism. Concurrency order: C1 → C4 → C8 → C16 → C32 2K context Prefill: 2,379 → 1,939 → 1,803 → 1,623 → 461 Decode: 48 → 106 → 104 → 100 → 104 32K context Prefill: 1,004 → 1,022 → 1,056 → 1,117 → 1,326…
mehr auf Arint.info
#agent #Agent #AGENT #cell #Commons #DeepSeek #GGUF #llama #LLAMA #Paris #Together #together #Unsloth #VLLM #vllm #vLLM #Wikimedia #arint_info
-
Quick Demonstration:
Translating text without using a web browser, no logins, using free tools:
#Linux
#LinuxMint
#AI
#llama
#ollama(This translation revealed something that the philosopher, Plato, said, which surprised me.... and his warning is a good one.)
-
Quick Demonstration:
Translating text without using a web browser, no logins, using free tools:
#Linux
#LinuxMint
#AI
#llama
#ollama(This translation revealed something that the philosopher, Plato, said, which surprised me.... and his warning is a good one.)
-
2/ Ich habe mir dann noch den Spaß gemacht, einen Kreis für #Llama hinzuzufügen. Warstadt & Bowerman haben darauf hingewiesen, dass der Input von GPT-3 20.000 Jahren Input eines Kindes wäre. Bei Llama wären das 1.500.000 Jahre. Janz schön lange, oder?
Das bedeutet, dass menschlicher #Sprachererwerb gaaanz anders abläuft.
-
2/ Ich habe mir dann noch den Spaß gemacht, einen Kreis für #Llama hinzuzufügen. Warstadt & Bowerman haben darauf hingewiesen, dass der Input von GPT-3 20.000 Jahren Input eines Kindes wäre. Bei Llama wären das 1.500.000 Jahre. Janz schön lange, oder?
Das bedeutet, dass menschlicher #Sprachererwerb gaaanz anders abläuft.
-
Install llama.cpp, run GGUF models with llama-cli, and serve OpenAI-compatible APIs using llama-server. Key flags, examples, and tuning tips with a short commands cheatsheet
#Cheatsheet #AI #LLM #DevOps #OpenAI #API #SelfHosting #Prometheus #llama.cpp
-
Install llama.cpp, run GGUF models with llama-cli, and serve OpenAI-compatible APIs using llama-server. Key flags, examples, and tuning tips with a short commands cheatsheet
#Cheatsheet #AI #LLM #DevOps #OpenAI #API #SelfHosting #Prometheus #llama.cpp
-
-
RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…
mehr auf Arint.info
#AGENT #Agent #agent #Apple #llama #nitter #Qwen36 #qwen36 #Wikipedia #arint_info
-
Ok, endlich #OpenWebUI und #llama.cpp und #ComfyUI zum editieren von Bildern überreden können. Dann auf ein schönes Wochenende.. fast. #ki #opensource #unabhängig