home.social

#qwen36 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #qwen36, aggregated by home.social.

fetched live
  1. 如果說智力的話 #MuseGlimmer 沒有超過 #Qwen36 27b 但它的確是會自己測試發現問題並修正,有點像AGI,配搭memory/skill 可以做到大部份 frontier models 的工作,用來省錢是不錯的

  2. RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…

    mehr auf Arint.info

    #AGENT #Agent #agent #Apple #llama #nitter #Qwen36 #qwen36 #Wikipedia #arint_info

    https://x.com/DataChaz/status/2080832572381393307#m

  3. RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…

    mehr auf Arint.info

    #AGENT #Agent #agent #Apple #llama #nitter #Qwen36 #qwen36 #Wikipedia #arint_info

    https://x.com/DataChaz/status/2080832572381393307#m

  4. RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…

    mehr auf Arint.info

    #AGENT #Agent #agent #Apple #llama #nitter #Qwen36 #qwen36 #Wikipedia #arint_info

    https://x.com/DataChaz/status/2080832572381393307#m

  5. RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…

    mehr auf Arint.info

    #AGENT #Agent #agent #Apple #llama #nitter #Qwen36 #qwen36 #Wikipedia #arint_info

    https://x.com/DataChaz/status/2080832572381393307#m

  6. RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…

    mehr auf Arint.info

    #Agent #AGENT #agent #Apple #llama #nitter #qwen36 #Qwen36 #Wikipedia #arint_info

    https://x.com/DataChaz/status/2080832572381393307#m

  7. RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…

    mehr auf Arint.info

    #Agent #AGENT #agent #Apple #llama #nitter #qwen36 #Qwen36 #Wikipedia #arint_info

    https://x.com/DataChaz/status/2080832572381393307#m

  8. RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…

    mehr auf Arint.info

    #Agent #AGENT #agent #Apple #llama #nitter #qwen36 #Qwen36 #Wikipedia #arint_info

    https://x.com/DataChaz/status/2080832572381393307#m

  9. RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…

    mehr auf Arint.info

    #Agent #AGENT #agent #Apple #llama #nitter #qwen36 #Qwen36 #Wikipedia #arint_info

    https://x.com/DataChaz/status/2080832572381393307#m

  10. Мини‑ПК на Strix Halo под параллельной нагрузкой: 236 tok/s на 32 одновременных запросах и три ошибки

    Один Beelink GTR9 Pro на Ryzen AI Max+ 395 выдал 236 tok/s суммарной генерации на 32 одновременных запросах в коротких прогонах и удержал в среднем 226 tok/s за 30 минут непрерывной нагрузки без тротлинга. По дороге к этим числам я нашёл воспроизводимый провал пропускной способности, который сначала выглядел свойством одной модели, увидел, как спекулятивный декодинг на моём стеке превращается из ускорителя в налог, и трижды чуть не опубликовал неверные выводы. Каждый раз спасали контрольные замеры, все три истории здесь, в статье. Харнессы, конфиги, полные таблицы, значения по каждому ключевому и финальному прогону, поминутные ряды выносливости и телеметрия лежат в открытом репозитории . Логи отдельных запросов харнесс не вёл, поэтому пересчитать можно всё до уровня прогона, но не глубже. Числа сняты на одном конкретном стенде; что из этого переносится на другие стеки, а что нет, оговорено по ходу текста.

    habr.com/ru/articles/1060520/

    #Strix_Halo #Ryzen_AI_Max #llamacpp #локальные_LLM #инференс #бенчмарки #Vulkan #Gemma_4 #Qwen36 #спекулятивный_декодинг

  11. Мини‑ПК на Strix Halo под параллельной нагрузкой: 236 tok/s на 32 одновременных запросах и три ошибки

    Один Beelink GTR9 Pro на Ryzen AI Max+ 395 выдал 236 tok/s суммарной генерации на 32 одновременных запросах в коротких прогонах и удержал в среднем 226 tok/s за 30 минут непрерывной нагрузки без тротлинга. По дороге к этим числам я нашёл воспроизводимый провал пропускной способности, который сначала выглядел свойством одной модели, увидел, как спекулятивный декодинг на моём стеке превращается из ускорителя в налог, и трижды чуть не опубликовал неверные выводы. Каждый раз спасали контрольные замеры, все три истории здесь, в статье. Харнессы, конфиги, полные таблицы, значения по каждому ключевому и финальному прогону, поминутные ряды выносливости и телеметрия лежат в открытом репозитории . Логи отдельных запросов харнесс не вёл, поэтому пересчитать можно всё до уровня прогона, но не глубже. Числа сняты на одном конкретном стенде; что из этого переносится на другие стеки, а что нет, оговорено по ходу текста.

    habr.com/ru/articles/1060520/

    #Strix_Halo #Ryzen_AI_Max #llamacpp #локальные_LLM #инференс #бенчмарки #Vulkan #Gemma_4 #Qwen36 #спекулятивный_декодинг

  12. Мини‑ПК на Strix Halo под параллельной нагрузкой: 236 tok/s на 32 одновременных запросах и три ошибки

    Один Beelink GTR9 Pro на Ryzen AI Max+ 395 выдал 236 tok/s суммарной генерации на 32 одновременных запросах в коротких прогонах и удержал в среднем 226 tok/s за 30 минут непрерывной нагрузки без тротлинга. По дороге к этим числам я нашёл воспроизводимый провал пропускной способности, который сначала выглядел свойством одной модели, увидел, как спекулятивный декодинг на моём стеке превращается из ускорителя в налог, и трижды чуть не опубликовал неверные выводы. Каждый раз спасали контрольные замеры, все три истории здесь, в статье. Харнессы, конфиги, полные таблицы, значения по каждому ключевому и финальному прогону, поминутные ряды выносливости и телеметрия лежат в открытом репозитории . Логи отдельных запросов харнесс не вёл, поэтому пересчитать можно всё до уровня прогона, но не глубже. Числа сняты на одном конкретном стенде; что из этого переносится на другие стеки, а что нет, оговорено по ходу текста.

    habr.com/ru/articles/1060520/

    #Strix_Halo #Ryzen_AI_Max #llamacpp #локальные_LLM #инференс #бенчмарки #Vulkan #Gemma_4 #Qwen36 #спекулятивный_декодинг

  13. ИИ Qwen3.6-27B запустили на смартфоне: 1 бит на вес и 90% интеллекта оригинала

    Стартап PrismML представил Bonsai 27B — сжатые версии открытой модели Qwen3.6-27B, младшая из которых стала первой нейросетью такого класса, которая помещается в память смартфона. Веса выложены на Hugging Face под лицензией Apache 2.0, а в демонстрациях PrismML модель работает прямо на iPhone 17 Pro Max — рассуждает, вызывает инструменты и разбирает скриншоты без единого обращения к облаку.

    habr.com/ru/articles/1059572/

    #qwen36 #iphone_17_pro_max #Bonsai_27B

  14. ИИ Qwen3.6-27B запустили на смартфоне: 1 бит на вес и 90% интеллекта оригинала

    Стартап PrismML представил Bonsai 27B — сжатые версии открытой модели Qwen3.6-27B, младшая из которых стала первой нейросетью такого класса, которая помещается в память смартфона. Веса выложены на Hugging Face под лицензией Apache 2.0, а в демонстрациях PrismML модель работает прямо на iPhone 17 Pro Max — рассуждает, вызывает инструменты и разбирает скриншоты без единого обращения к облаку.

    habr.com/ru/articles/1059572/

    #qwen36 #iphone_17_pro_max #Bonsai_27B

  15. ИИ Qwen3.6-27B запустили на смартфоне: 1 бит на вес и 90% интеллекта оригинала

    Стартап PrismML представил Bonsai 27B — сжатые версии открытой модели Qwen3.6-27B, младшая из которых стала первой нейросетью такого класса, которая помещается в память смартфона. Веса выложены на Hugging Face под лицензией Apache 2.0, а в демонстрациях PrismML модель работает прямо на iPhone 17 Pro Max — рассуждает, вызывает инструменты и разбирает скриншоты без единого обращения к облаку.

    habr.com/ru/articles/1059572/

    #qwen36 #iphone_17_pro_max #Bonsai_27B

  16. RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x

    mehr auf Arint.info

    #qwen36 #arint_info

    https://x.com/superalesha/status/2076356044104585667#m

  17. RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x

    mehr auf Arint.info

    #qwen36 #arint_info

    https://x.com/superalesha/status/2076356044104585667#m

  18. RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x

    mehr auf Arint.info

    #qwen36 #arint_info

    https://x.com/superalesha/status/2076356044104585667#m

  19. RT @spiritbuun: My previous best decode on a single 3090 with Qwen3.6 27B was 206 tok/s. Today I beat it. 219 tok/s. One 3090. DFlash perf work has been PUSHED, get it now on buun-llama. Several more DFlash improvements are in development- expect several more incremental improvements this week

    mehr auf Arint.info

    #llama #Qwen36 #arint_info

    https://x.com/spiritbuun/status/2076737488681349325#m

  20. RT @spiritbuun: My previous best decode on a single 3090 with Qwen3.6 27B was 206 tok/s. Today I beat it. 219 tok/s. One 3090. DFlash perf work has been PUSHED, get it now on buun-llama. Several more DFlash improvements are in development- expect several more incremental improvements this week

    mehr auf Arint.info

    #llama #Qwen36 #arint_info

    https://x.com/spiritbuun/status/2076737488681349325#m

  21. RT @spiritbuun: My previous best decode on a single 3090 with Qwen3.6 27B was 206 tok/s. Today I beat it. 219 tok/s. One 3090. DFlash perf work has been PUSHED, get it now on buun-llama. Several more DFlash improvements are in development- expect several more incremental improvements this week

    mehr auf Arint.info

    #llama #Qwen36 #arint_info

    https://x.com/spiritbuun/status/2076737488681349325#m

  22. RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x

    mehr auf Arint.info

    #qwen36 #arint_info

    https://x.com/superalesha/status/2076356044104585667#m

  23. RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x

    mehr auf Arint.info

    #qwen36 #arint_info

    https://x.com/superalesha/status/2076356044104585667#m

  24. RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x

    mehr auf Arint.info

    #qwen36 #arint_info

    https://x.com/superalesha/status/2076356044104585667#m

  25. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  26. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  27. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  28. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  29. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  30. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  31. RT @MiaAI_lab: Welches lokale Modell ist das beste für Agentic Workflows auf einem einzelnen @NVIDIAAI DGX Spark? (oder einem anderen Rig mit 96-128 GB VRAM) Nach der Auswertung von 84 Szenarien, 16 Kategorien und jeweils 8 Durchläufen im Hermes-Agent-Stil für die mehrstufige Tool-Orchestrierung gibt es einen sehr klaren Sieger. 🏆 Qwen 3.6 35B A3B Q8KXL liegt auf Platz 1. Es ist das einzige Modell, das durchweg perfekte Ergebnisse erzielte und keine katastrophalen Ausfälle aufwies. Die vollständige Rangliste: Qwen 3.6 35B A3B UD Q8KXL — 91,0 | Qwen 3.6 27B NVFP4 — 89,0 | Qwopus 3.6 27B Coder MTP — 85,2 | DeepSeek V4 Flash Q2 — 86,5 | Agents-A1 Q80 — 83,4 | Gemma 4 26B — 81,4 | Nemotron 3 Nano Omni 30B — 79,0 Fazit: Wenn du 2026 Agenten lokal auf einem DGX Spark oder einem anderen Rig mit 96–128 GB VRAM betreibst, ist Qwen 3.6 35B Q8KXL derzeit die beste Wahl. Vollständiger Bericht + tiefgehende Analyse 👇 github.com/MiaAI-Lab/Best-Loca

    mehr auf Arint.info

    #AgenticAI #AIAgents #LocalLLM #MachineLearning #NVIDIADGX #Qwen36 #arint_info

    https://x.com/MiaAI_lab/status/2074093556545749235#m

  32. RT @MiaAI_lab: Welches lokale Modell ist das beste für Agentic Workflows auf einem einzelnen @NVIDIAAI DGX Spark? (oder einem anderen Rig mit 96-128 GB VRAM) Nach der Auswertung von 84 Szenarien, 16 Kategorien und jeweils 8 Durchläufen im Hermes-Agent-Stil für die mehrstufige Tool-Orchestrierung gibt es einen sehr klaren Sieger. 🏆 Qwen 3.6 35B A3B Q8KXL liegt auf Platz 1. Es ist das einzige Modell, das durchweg perfekte Ergebnisse erzielte und keine katastrophalen Ausfälle aufwies. Die vollständige Rangliste: Qwen 3.6 35B A3B UD Q8KXL — 91,0 | Qwen 3.6 27B NVFP4 — 89,0 | Qwopus 3.6 27B Coder MTP — 85,2 | DeepSeek V4 Flash Q2 — 86,5 | Agents-A1 Q80 — 83,4 | Gemma 4 26B — 81,4 | Nemotron 3 Nano Omni 30B — 79,0 Fazit: Wenn du 2026 Agenten lokal auf einem DGX Spark oder einem anderen Rig mit 96–128 GB VRAM betreibst, ist Qwen 3.6 35B Q8KXL derzeit die beste Wahl. Vollständiger Bericht + tiefgehende Analyse 👇 github.com/MiaAI-Lab/Best-Loca

    mehr auf Arint.info

    #AgenticAI #AIAgents #LocalLLM #MachineLearning #NVIDIADGX #Qwen36 #arint_info

    https://x.com/MiaAI_lab/status/2074093556545749235#m

  33. RT @MiaAI_lab: Welches lokale Modell ist das beste für Agentic Workflows auf einem einzelnen @NVIDIAAI DGX Spark? (oder einem anderen Rig mit 96-128 GB VRAM) Nach der Auswertung von 84 Szenarien, 16 Kategorien und jeweils 8 Durchläufen im Hermes-Agent-Stil für die mehrstufige Tool-Orchestrierung gibt es einen sehr klaren Sieger. 🏆 Qwen 3.6 35B A3B Q8KXL liegt auf Platz 1. Es ist das einzige Modell, das durchweg perfekte Ergebnisse erzielte und keine katastrophalen Ausfälle aufwies. Die vollständige Rangliste: Qwen 3.6 35B A3B UD Q8KXL — 91,0 | Qwen 3.6 27B NVFP4 — 89,0 | Qwopus 3.6 27B Coder MTP — 85,2 | DeepSeek V4 Flash Q2 — 86,5 | Agents-A1 Q80 — 83,4 | Gemma 4 26B — 81,4 | Nemotron 3 Nano Omni 30B — 79,0 Fazit: Wenn du 2026 Agenten lokal auf einem DGX Spark oder einem anderen Rig mit 96–128 GB VRAM betreibst, ist Qwen 3.6 35B Q8KXL derzeit die beste Wahl. Vollständiger Bericht + tiefgehende Analyse 👇 github.com/MiaAI-Lab/Best-Loca

    mehr auf Arint.info

    #AgenticAI #AIAgents #LocalLLM #MachineLearning #NVIDIADGX #Qwen36 #arint_info

    https://x.com/MiaAI_lab/status/2074093556545749235#m

  34. RT @Tono_Ken3: Und bei StrixHalo's DearfStar4 läuft DeepSeek-V4-Flash mit 16 TPS. Da der Prefill den KV-Cache auf einer Optane-SSD speichert, ist die Geschwindigkeit wirklich beeindruckend. Es könnte auch gut sein, diesen Hermes als Sub-Agent von Qwen3.6's Lnagent aufzurufen. Das lokale Agenten-System besteht aus diesen drei Modellen: Qwen3.6-35b-a3b-nvfp4, DeepSeek-V4Flash-IQ2 und GLM-5.2-UQ4. Es ist übersichtlich. TonoKen3🤖Local-LLM&Robot🏁とのけん3 (@TonoKen3) Ja genau. Die Möglichkeit, GLM-5.2 lokal einzusetzen, schafft ein Gefühl von innerem Frieden. Für 90% der Fälle reicht die schnelle Antwort von Qwen3.6-35b. 130 TPS bieten eine komfortable Reaktionsgeschwindigkeit, die sogar die Nutzung geschlossener Modelle übertrifft. Bei der Inferenz verbraucht das System 550W, im Standby nur 200W. Das ist genau das, wonach man sucht. — nitter.net/TonoKen3/status/207

    mehr auf Arint.info

    #AIInfrastructure #DeepSeekV4 #GLM52 #LocalLLM #Qwen36 #TonoKen3 #arint_info

    https://x.com/Tono_Ken3/status/2073898742496047515#m

  35. RT @Tono_Ken3: Und bei StrixHalo's DearfStar4 läuft DeepSeek-V4-Flash mit 16 TPS. Da der Prefill den KV-Cache auf einer Optane-SSD speichert, ist die Geschwindigkeit wirklich beeindruckend. Es könnte auch gut sein, diesen Hermes als Sub-Agent von Qwen3.6's Lnagent aufzurufen. Das lokale Agenten-System besteht aus diesen drei Modellen: Qwen3.6-35b-a3b-nvfp4, DeepSeek-V4Flash-IQ2 und GLM-5.2-UQ4. Es ist übersichtlich. TonoKen3🤖Local-LLM&Robot🏁とのけん3 (@TonoKen3) Ja genau. Die Möglichkeit, GLM-5.2 lokal einzusetzen, schafft ein Gefühl von innerem Frieden. Für 90% der Fälle reicht die schnelle Antwort von Qwen3.6-35b. 130 TPS bieten eine komfortable Reaktionsgeschwindigkeit, die sogar die Nutzung geschlossener Modelle übertrifft. Bei der Inferenz verbraucht das System 550W, im Standby nur 200W. Das ist genau das, wonach man sucht. — nitter.net/TonoKen3/status/207

    mehr auf Arint.info

    #AIInfrastructure #DeepSeekV4 #GLM52 #LocalLLM #Qwen36 #TonoKen3 #arint_info

    https://x.com/Tono_Ken3/status/2073898742496047515#m

  36. RT @Tono_Ken3: Und bei StrixHalo's DearfStar4 läuft DeepSeek-V4-Flash mit 16 TPS. Da der Prefill den KV-Cache auf einer Optane-SSD speichert, ist die Geschwindigkeit wirklich beeindruckend. Es könnte auch gut sein, diesen Hermes als Sub-Agent von Qwen3.6's Lnagent aufzurufen. Das lokale Agenten-System besteht aus diesen drei Modellen: Qwen3.6-35b-a3b-nvfp4, DeepSeek-V4Flash-IQ2 und GLM-5.2-UQ4. Es ist übersichtlich. TonoKen3🤖Local-LLM&Robot🏁とのけん3 (@TonoKen3) Ja genau. Die Möglichkeit, GLM-5.2 lokal einzusetzen, schafft ein Gefühl von innerem Frieden. Für 90% der Fälle reicht die schnelle Antwort von Qwen3.6-35b. 130 TPS bieten eine komfortable Reaktionsgeschwindigkeit, die sogar die Nutzung geschlossener Modelle übertrifft. Bei der Inferenz verbraucht das System 550W, im Standby nur 200W. Das ist genau das, wonach man sucht. — nitter.net/TonoKen3/status/207

    mehr auf Arint.info

    #AIInfrastructure #DeepSeekV4 #GLM52 #LocalLLM #Qwen36 #TonoKen3 #arint_info

    https://x.com/Tono_Ken3/status/2073898742496047515#m

  37. Контекстная инженерия для слабой локальной модели: как мы делаем среднюю модель надёжной

    Принято думать, что качество ИИ-агента упирается в размер модели. Но когда модель работает локально, в закрытом контуре и на ограниченном железе, брать «побольше» особо некуда. И оказывается, что главный рычаг не модель, а контекст: что вы ей показываете, в каком порядке и как фильтруете. Причём «контекст» здесь — это сборка под то, кто спрашивает, откуда и о чём, плюс честный порог релевантности и продуманный порядок секций. На сильной облачной модели небрежный контекст прощается запасом по reasoning; на средней локальной — нет. Об этом и статья.

    habr.com/ru/companies/1forma/a

    #llm #onpremise #qwen36 #ai #aiагенты #искусственный_интеллект #автоматизация_процессов #корпоративные_системы #lowcode #bpms

  38. Контекстная инженерия для слабой локальной модели: как мы делаем среднюю модель надёжной

    Принято думать, что качество ИИ-агента упирается в размер модели. Но когда модель работает локально, в закрытом контуре и на ограниченном железе, брать «побольше» особо некуда. И оказывается, что главный рычаг не модель, а контекст: что вы ей показываете, в каком порядке и как фильтруете. Причём «контекст» здесь — это сборка под то, кто спрашивает, откуда и о чём, плюс честный порог релевантности и продуманный порядок секций. На сильной облачной модели небрежный контекст прощается запасом по reasoning; на средней локальной — нет. Об этом и статья.

    habr.com/ru/companies/1forma/a

    #llm #onpremise #qwen36 #ai #aiагенты #искусственный_интеллект #автоматизация_процессов #корпоративные_системы #lowcode #bpms

  39. Контекстная инженерия для слабой локальной модели: как мы делаем среднюю модель надёжной

    Принято думать, что качество ИИ-агента упирается в размер модели. Но когда модель работает локально, в закрытом контуре и на ограниченном железе, брать «побольше» особо некуда. И оказывается, что главный рычаг не модель, а контекст: что вы ей показываете, в каком порядке и как фильтруете. Причём «контекст» здесь — это сборка под то, кто спрашивает, откуда и о чём, плюс честный порог релевантности и продуманный порядок секций. На сильной облачной модели небрежный контекст прощается запасом по reasoning; на средней локальной — нет. Об этом и статья.

    habr.com/ru/companies/1forma/a

    #llm #onpremise #qwen36 #ai #aiагенты #искусственный_интеллект #автоматизация_процессов #корпоративные_системы #lowcode #bpms

  40. So I saw this Spider-man reference over on Reddit, and I realized I wasn't familiar with what Peter is referencing here.

    en.wikipedia.org/wiki/Brachist

    It's pretty neat.

    But I also had a side thought: let's throw it at #Qwen36 and have it make an interactive demonstration?

    After about ~15 mins of churn, it made a single file HTML: scratch.network47.org/s/a8yocv

    This is WITHOUT using Wikipedia as a reference. Purely from the model.

    #llm #localllm

  41. So I saw this Spider-man reference over on Reddit, and I realized I wasn't familiar with what Peter is referencing here.

    en.wikipedia.org/wiki/Brachist

    It's pretty neat.

    But I also had a side thought: let's throw it at #Qwen36 and have it make an interactive demonstration?

    After about ~15 mins of churn, it made a single file HTML: scratch.network47.org/s/a8yocv

    This is WITHOUT using Wikipedia as a reference. Purely from the model.

    #llm #localllm

  42. So I saw this Spider-man reference over on Reddit, and I realized I wasn't familiar with what Peter is referencing here.

    en.wikipedia.org/wiki/Brachist

    It's pretty neat.

    But I also had a side thought: let's throw it at #Qwen36 and have it make an interactive demonstration?

    After about ~15 mins of churn, it made a single file HTML: scratch.network47.org/s/a8yocv

    This is WITHOUT using Wikipedia as a reference. Purely from the model.

    #llm #localllm

  43. So I saw this Spider-man reference over on Reddit, and I realized I wasn't familiar with what Peter is referencing here.

    en.wikipedia.org/wiki/Brachist

    It's pretty neat.

    But I also had a side thought: let's throw it at #Qwen36 and have it make an interactive demonstration?

    After about ~15 mins of churn, it made a single file HTML: scratch.network47.org/s/a8yocv

    This is WITHOUT using Wikipedia as a reference. Purely from the model.

    #llm #localllm

  44. So I saw this Spider-man reference over on Reddit, and I realized I wasn't familiar with what Peter is referencing here.

    en.wikipedia.org/wiki/Brachist

    It's pretty neat.

    But I also had a side thought: let's throw it at #Qwen36 and have it make an interactive demonstration?

    After about ~15 mins of churn, it made a single file HTML: scratch.network47.org/s/a8yocv

    This is WITHOUT using Wikipedia as a reference. Purely from the model.

    #llm #localllm

  45. Тесты бюджетных сборок для ИИ до 100к рублей

    Локальный ИИ не должен стоить как автомобиль. Мне стало интересно: возможен ли жизнеспособный инференс на CPU и что реально дают дешевые GPU (вроде Tesla V100 или CMP 40HX). Я собрал несколько бюджетных конфигураций до 100к, потестил актуальные модели и попытался понять, что важнее для скорости: канальность памяти или частота. Сравнил дешевые AM4 и Threadripper, замерил токены в секунду и построил графики. Делюсь результатами.

    habr.com/ru/articles/1053118/

    #ai #ии #gpt #selfhosted #gpu #cpu #llamacpp #qwen36 #gemma4

  46. Тесты бюджетных сборок для ИИ до 100к рублей

    Локальный ИИ не должен стоить как автомобиль. Мне стало интересно: возможен ли жизнеспособный инференс на CPU и что реально дают дешевые GPU (вроде Tesla V100 или CMP 40HX). Я собрал несколько бюджетных конфигураций до 100к, потестил актуальные модели и попытался понять, что важнее для скорости: канальность памяти или частота. Сравнил дешевые AM4 и Threadripper, замерил токены в секунду и построил графики. Делюсь результатами.

    habr.com/ru/articles/1053118/

    #ai #ии #gpt #selfhosted #gpu #cpu #llamacpp #qwen36 #gemma4

  47. Тесты бюджетных сборок для ИИ до 100к рублей

    Локальный ИИ не должен стоить как автомобиль. Мне стало интересно: возможен ли жизнеспособный инференс на CPU и что реально дают дешевые GPU (вроде Tesla V100 или CMP 40HX). Я собрал несколько бюджетных конфигураций до 100к, потестил актуальные модели и попытался понять, что важнее для скорости: канальность памяти или частота. Сравнил дешевые AM4 и Threadripper, замерил токены в секунду и построил графики. Делюсь результатами.

    habr.com/ru/articles/1053118/

    #ai #ии #gpt #selfhosted #gpu #cpu #llamacpp #qwen36 #gemma4

  48. 之前用 #Qwen36 Plus 都試過可以自動化做網頁應用的開發,它會自己開啟Google Chrome,通過 Chrome devtools MCP,因為有視覺能力所以都可以自己做開發和研究,唯一的問題是GUI有太多坑是AI模仿不到的,例如我之前失敗的是打中文字輸入法有問題,還有無故會失去打字的焦點

  49. Локальные LLM на Arch Linux и как увеличить скорость генерации в 20 раз

    Приветствую всех читателей Хабра, в этой статье я хочу поделиться своим опытом в запуске локальных LLM, протестировать работоспособность интересных моделей на своем железе, рассказать, как я увеличил скорость генерации на одной из нейросетей в 20 раз (я не преувеличиваю). Но об этом чуть позже, а начну я повествование с описания своего железа.

    habr.com/ru/articles/1045898/

    #arch_linux #llamacpp #ollama #qwen36 #gemma4 #github #huggingface #intel_arc_b580

  50. Локальные LLM на Arch Linux и как увеличить скорость генерации в 20 раз

    Приветствую всех читателей Хабра, в этой статье я хочу поделиться своим опытом в запуске локальных LLM, протестировать работоспособность интересных моделей на своем железе, рассказать, как я увеличил скорость генерации на одной из нейросетей в 20 раз (я не преувеличиваю). Но об этом чуть позже, а начну я повествование с описания своего железа.

    habr.com/ru/articles/1045898/

    #arch_linux #llamacpp #ollama #qwen36 #gemma4 #github #huggingface #intel_arc_b580