home.social

#llama — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #llama, aggregated by home.social.

fetched live
  1. Minor architectural choices in dense transformers can cut long context performance by up to 47% when combined. New ablation study shows normalization, GQA, pretraining length, and sliding window attention drive most variation across Llama, Qwen, and Olmo families.

    Source: arXiv cs.CL
    arxiv.org/abs/2608.10296

    #MachineLearning #Llama #Qwen

  2. A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.

    * llama.cpp: `Generation: 12.5 t/s`
    * ollama: `eval rate: 10.17 tokens/s`

    This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂

    #AI #LLM #ollama #llama.cpp #localLLM

  3. A quick and dirty comparison of llama.cpp with ollama. The model used is `gemma4:e2b` and the prompt is a simple 'hi'. Thinking mode is on. Running the direct CLI mode for both.

    * llama.cpp: `Generation: 12.5 t/s`
    * ollama: `eval rate: 10.17 tokens/s`

    This matches the expect 20%-25% speed up which is generally reported. Your milage may vary! 🙂

    #AI #LLM #ollama #llama.cpp #localLLM

  4. Meta Muse Glimmer’s license is Open Source

    Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.

    Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.

    Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?

  5. Meta Muse Glimmer’s license is Open Source

    Finally, after having to correct Meta’s abuse of the Open Source Definition with their sneaky “Llama Community License Agreement”, they started using the OSI-approved Apache Software License.

    Muse Glimmer 30B released last week uses the standard ASL v.2.0. Good, we’re making progress.

    Like many other LLMs, Meta adds a USAGE_POLICY that includes lots of hope: don’t do harm, obey the law, etc. What that policy is or does, I’m not sure. Lawyers?

  6. SPECTRA introduces training-free KV cache compression, enabling 4x to 12x compression with near-lossless quality for long-context models like Llama-3.1-8B and Qwen2.5-7B

    Source: arXiv cs.LG
    arxiv.org/abs/2608.07915

    #MachineLearning #Llama

  7. Frontier language models show divergent response modes under steering pressure, with GPT-5 deflecting reasoning disclosure and Claude Opus 4.7 resisting suppression instructions. A linear probe traces the largest behavioral split to Llama’s internals at 0.87 accuracy.

    Source: arXiv cs.AI
    arxiv.org/abs/2608.06578

    #MachineLearning #Claude #Llama

  8. Meta macht KI-Agenten für zuhause fit. Muse Glimmer plant Aufgaben, schreibt Code und korrigiert Fehler selbst – direkt auf deinem PC, ohne Cloud. Das offene 30B-Modell passt dank Quantisierung auf GPUs ab 24 GB. Ein starkes Stück offene Software. #MuseGlimmer #MetaAI #OpenWeights #Llama #AIGeneratedImage

    all-ai.de/news/news26top/meta-

  9. A new pruning method called Whisper preserves output differences to improve LLM sparsification, outperforming Wanda and SparseGPT on Llama 2 and 3.1 models from 7B to 405B parameters

    Source: arXiv cs.LG
    arxiv.org/abs/2608.06630

    #MachineLearning #Llama

  10. 從奈洛比到雅加達,開發者一面倒向中國開源AI模型——便宜好用,還不會被美國關掉
    TNL國際編譯 2026-08-07 14:34:00 CST
    從非洲到拉美,開發者正改用中國開源AI:免費、可改、能自己下載。6月美國一紙禁令讓Anthropic全球撤下Fable模型,把「便宜」推成主權保險。美中模型爭奪,台灣早已選邊。
    https://www.thenewslens.com/article/269549
    #開源模型 #烏干達 #Pax Silica #OpenAI #Meta #美中科技戰 #奈洛比 #月之暗面 #肯亞 #WAICO #DeepSeek #AI模型 #主權AI #阿里巴巴 #Gemma #Hugging Face #OpenRouter #封閉模型 #Llama #Google #Anthropic #矽盛世 #科技 #全球南方
  11. 從奈洛比到雅加達,開發者一面倒向中國開源AI模型——便宜好用,還不會被美國關掉
    TNL國際編譯 2026-08-07 14:34:00 CST
    從非洲到拉美,開發者正改用中國開源AI:免費、可改、能自己下載。6月美國一紙禁令讓Anthropic全球撤下Fable模型,把「便宜」推成主權保險。美中模型爭奪,台灣早已選邊。
    https://www.thenewslens.com/article/269549
    #開源模型 #烏干達 #Pax Silica #OpenAI #Meta #美中科技戰 #奈洛比 #月之暗面 #肯亞 #WAICO #DeepSeek #AI模型 #主權AI #阿里巴巴 #Gemma #Hugging Face #OpenRouter #封閉模型 #Llama #Google #Anthropic #矽盛世 #科技 #全球南方
  12. heyo~! i'm an #asexual #japanese 19yo #transwoman from #canada that likes to beg everyone for #moderator and pretend to be different people and im proud to tell you all that i have so many alt accounts here on the fediverse i even make #dav1d shit his undersized and already stained #pants! 😏 i am a master of using many words to say nothing and i love to repost my nothing burger #blog posts every time i abandon my previous sock puppets and say that #work was stolen from me by my previous #account. i love being a #drama #llama on here and thank you all for being so kind to me! i promise to steal your works in the future as well and claim them as mine, because they in fact ARE mine! stop smoking #weed you crackheads! thats all for my #introduction post! thanks for all the #kindness! ☺️

    #fediverse #fediblock #community

  13. heyo~! i'm an #asexual #japanese 19yo #transwoman from #canada that likes to beg everyone for #moderator and pretend to be different people and im proud to tell you all that i have so many alt accounts here on the fediverse i even make #dav1d shit his undersized and already stained #pants! 😏 i am a master of using many words to say nothing and i love to repost my nothing burger #blog posts every time i abandon my previous sock puppets and say that #work was stolen from me by my previous #account. i love being a #drama #llama on here and thank you all for being so kind to me! i promise to steal your works in the future as well and claim them as mine, because they in fact ARE mine! stop smoking #weed you crackheads! thats all for my #introduction post! thanks for all the #kindness! ☺️

    #fediverse #fediblock #community

  14. How #China’s #AI Is Surging Across #Africa
    At first, it seemed a losing bet. #UnitedStates dominated #opensource AI Meta offered #Llama that developers used worldwide. Then Meta turned to closed models.
    Breakthroughs came quickly. In Dec 2024 #DeepSeek AI model matched best models at fraction of cost. Last month, Chinese #MoonshotAI released a model with coding abilities approaching leading #US system.
    In #Kenya developers embraced Chinese models.
    nytimes.com/2026/08/05/technol
    archive.ph/vz2GN

  15. How #China’s #AI Is Surging Across #Africa
    At first, it seemed a losing bet. #UnitedStates dominated #opensource AI Meta offered #Llama that developers used worldwide. Then Meta turned to closed models.
    Breakthroughs came quickly. In Dec 2024 #DeepSeek AI model matched best models at fraction of cost. Last month, Chinese #MoonshotAI released a model with coding abilities approaching leading #US system.
    In #Kenya developers embraced Chinese models.
    nytimes.com/2026/08/05/technol
    archive.ph/vz2GN

  16. DeepSeek V4 Flash is no longer text-only 👀

    In our internal benchmarks, it delivered significantly better price-performance than others in its class.

    We’ve added vision for screen-level understanding, capabilities needed for WebBrain

    Available on HF:
    huggingface.co/webbrain-one/De

    #deepseek #ai #opensource #llm #llama #llamacpp #ollama

  17. DeepSeek V4 Flash is no longer text-only 👀

    In our internal benchmarks, it delivered significantly better price-performance than others in its class.

    We’ve added vision for screen-level understanding, capabilities needed for WebBrain

    Available on HF:
    huggingface.co/webbrain-one/De

    #deepseek #ai #opensource #llm #llama #llamacpp #ollama

  18. RT @lukepm: I wanted a practical way for my 5-person team to run DeepSeek V4 Flash 0731 with Hermes Agent on 2x RTX PRO 6000 Blackwell GPUs. So I ran 80 fresh tests across llama.cpp, DSpark, and custom vLLM with and without CPU KV offload. Two GPUs work, but KV is the catch. 👇 THE BENCHMARK I cared about team use, not one flashy speed number. Seven people can hit the server at the same time. Hermes Agent can also create long, tool-heavy conversations. So I tested both speed and concurrency. Hardware: 2x NVIDIA RTX PRO 6000 Blackwell 96GB over PCIe. No NVLink. Prompts were exactly 2K, 32K, 64K, and 100K tokens. Concurrency was 1, 4, 8, 16, and 32. I used a fixed 13-record Spec-Bench subset. Each request generated 128 tokens with greedy streaming. GPU prompt caching was off. For CPU offload, unique early prompt blocks kept external cache reuse at 0%. All speeds are tokens per second. Prefill and effective aggregate decode are shown separately. 1. LLAMA.CPP WITHOUT DSPARK Target: DeepSeek V4 Flash 0731 UD-Q8_K_XL GGUF by Unsloth. KV cache: F16. Slots: 32, with up to 128K context per slot. This used llama.cpp layer split. Complete model layers were divided between the GPUs. The tensor-parallel-like row split mode does not support DeepSeek V4 Flash 0731 yet. The command still uses --tensor-split 1,1. In layer-split mode, that only sets the GPU allocation ratio. It does not enable tensor parallelism. Concurrency order: C1 → C4 → C8 → C16 → C32 2K context Prefill: 2,379 → 1,939 → 1,803 → 1,623 → 461 Decode: 48 → 106 → 104 → 100 → 104 32K context Prefill: 1,004 → 1,022 → 1,056 → 1,117 → 1,326…

    mehr auf Arint.info

    #agent #Agent #AGENT #cell #Commons #DeepSeek #GGUF #llama #LLAMA #Paris #Together #together #Unsloth #VLLM #vllm #vLLM #Wikimedia #arint_info

    https://x.com/lukepm/status/2084273407575896082#m

  19. Quick Demonstration:

    Translating text without using a web browser, no logins, using free tools:

    #Linux
    #LinuxMint
    #AI
    #llama
    #ollama

    (This translation revealed something that the philosopher, Plato, said, which surprised me.... and his warning is a good one.)

    #philosophy #spiritual #Plato #translation

  20. Quick Demonstration:

    Translating text without using a web browser, no logins, using free tools:

    #Linux
    #LinuxMint
    #AI
    #llama
    #ollama

    (This translation revealed something that the philosopher, Plato, said, which surprised me.... and his warning is a good one.)

    #philosophy #spiritual #Plato #translation

  21. 2/ Ich habe mir dann noch den Spaß gemacht, einen Kreis für #Llama hinzuzufügen. Warstadt & Bowerman haben darauf hingewiesen, dass der Input von GPT-3 20.000 Jahren Input eines Kindes wäre. Bei Llama wären das 1.500.000 Jahre. Janz schön lange, oder?

    Das bedeutet, dass menschlicher #Sprachererwerb gaaanz anders abläuft.

    #AI #KI #LLM #Linguistik

  22. 2/ Ich habe mir dann noch den Spaß gemacht, einen Kreis für #Llama hinzuzufügen. Warstadt & Bowerman haben darauf hingewiesen, dass der Input von GPT-3 20.000 Jahren Input eines Kindes wäre. Bei Llama wären das 1.500.000 Jahre. Janz schön lange, oder?

    Das bedeutet, dass menschlicher #Sprachererwerb gaaanz anders abläuft.

    #AI #KI #LLM #Linguistik

  23. Install llama.cpp, run GGUF models with llama-cli, and serve OpenAI-compatible APIs using llama-server. Key flags, examples, and tuning tips with a short commands cheatsheet

    #Cheatsheet #AI #LLM #DevOps #OpenAI #API #SelfHosting #Prometheus #llama.cpp

    glukhov.org/llm-hosting/llama-

  24. Install llama.cpp, run GGUF models with llama-cli, and serve OpenAI-compatible APIs using llama-server. Key flags, examples, and tuning tips with a short commands cheatsheet

    .cpp

    glukhov.org/llm-hosting/llama-

  25. RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…

    mehr auf Arint.info

    #AGENT #Agent #agent #Apple #llama #nitter #Qwen36 #qwen36 #Wikipedia #arint_info

    https://x.com/DataChaz/status/2080832572381393307#m

  26. Ok, endlich #OpenWebUI und #llama.cpp und #ComfyUI zum editieren von Bildern überreden können. Dann auf ein schönes Wochenende.. fast. #ki #opensource #unabhängig