#qwen36 — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #qwen36, aggregated by home.social.
-
Сколько инструментов показывать языковой модели: гипотеза, которую перечеркнул замер
У нашего ИИ-оркестратора более девяноста инструментов: работа с таблицами, документами, почтой, задачами, веб-страницами, презентациями и т.д. При обработке запросов модель на каждом шаге выбирает, какой из них вызвать. Вопрос, который периодически всплывает: сколько из них ей показывать за раз — все или только те, что подходят к запросу? Наш оркестратор умеет работать как с большими облачными моделями, так и с маленькими, которые помещаются на бытовой видеокарте. Минимальный рабочий вариант: RTX 3090 24Gb + Qwen3.6 27B. Цель - заставить минимальный рабочий вариант выполнять корпоративные задачи с приемлемым качеством и с приемлемой скоростью. Ясно что, для малой и большой моделей способ выбора инструментов должен отличаться, но как найти оптимальный вариант вопрос сложный. Обычно оркестратор показывает модели весь список инструментов и она уже решает какой нужно выбрать для ответа на запрос. Так делают Hermes, OpenClaw и Ouroboros. Ни один из них не выбирает инструменты по тексту запроса. У некоторых для слабых моделей и маленького контекста есть режим, где модели дают только каталог — имя и короткое описание — и инструмент «подключить», так что нужное она находит и загружает сама. Эти агенты рассчитаны на фронтир-модели уровня Claude и GPT-5, которые держат сотню инструментов без потерь. Моя задача другая: работать на локальной 27B в закрытом контуре. Для неё узкий набор измеримо лучше. Как устроен узкий набор Схема простая. Перед основным вызовом модели работает роутер: по тексту запроса он выбирает раздел — «датасеты», «почта», «презентации» ... — и внутри раздела подраздел. Модель получает недевяносто описаний инструментов, а около двадцати. О существовании остальных на этом шаге она не знает.
https://habr.com/ru/articles/1089742/
#ииагенты #ииассистент #llm #llmагент #оркестрация_агентов #оркестраторы #rtx_3090 #rtx_5090 #qwen36
-
Сколько инструментов показывать языковой модели: гипотеза, которую перечеркнул замер
У нашего ИИ-оркестратора более девяноста инструментов: работа с таблицами, документами, почтой, задачами, веб-страницами, презентациями и т.д. При обработке запросов модель на каждом шаге выбирает, какой из них вызвать. Вопрос, который периодически всплывает: сколько из них ей показывать за раз — все или только те, что подходят к запросу? Наш оркестратор умеет работать как с большими облачными моделями, так и с маленькими, которые помещаются на бытовой видеокарте. Минимальный рабочий вариант: RTX 3090 24Gb + Qwen3.6 27B. Цель - заставить минимальный рабочий вариант выполнять корпоративные задачи с приемлемым качеством и с приемлемой скоростью. Ясно что, для малой и большой моделей способ выбора инструментов должен отличаться, но как найти оптимальный вариант вопрос сложный. Обычно оркестратор показывает модели весь список инструментов и она уже решает какой нужно выбрать для ответа на запрос. Так делают Hermes, OpenClaw и Ouroboros. Ни один из них не выбирает инструменты по тексту запроса. У некоторых для слабых моделей и маленького контекста есть режим, где модели дают только каталог — имя и короткое описание — и инструмент «подключить», так что нужное она находит и загружает сама. Эти агенты рассчитаны на фронтир-модели уровня Claude и GPT-5, которые держат сотню инструментов без потерь. Моя задача другая: работать на локальной 27B в закрытом контуре. Для неё узкий набор измеримо лучше. Как устроен узкий набор Схема простая. Перед основным вызовом модели работает роутер: по тексту запроса он выбирает раздел — «датасеты», «почта», «презентации» ... — и внутри раздела подраздел. Модель получает недевяносто описаний инструментов, а около двадцати. О существовании остальных на этом шаге она не знает.
https://habr.com/ru/articles/1089742/
#ииагенты #ииассистент #llm #llmагент #оркестрация_агентов #оркестраторы #rtx_3090 #rtx_5090 #qwen36
-
35B-модель на RTX 5080 и RTX 3060: 109 ток/с через PCIe Gen2 x4
На материнке B450-Plus с урезанным вторым слотом PCIe Gen2 x4 Qwen3.6 выдала 109 ток/с. MoE обнуляет аргумент «одна видеокарта лучше двух».
https://habr.com/ru/articles/1088316/
#llamacpp #локальные_llm #moe #gguf #инференс_llm #vram_использование #pcie #qwen36 #Qwen_36 #локальный_ии
-
35B-модель на RTX 5080 и RTX 3060: 109 ток/с через PCIe Gen2 x4
На материнке B450-Plus с урезанным вторым слотом PCIe Gen2 x4 Qwen3.6 выдала 109 ток/с. MoE обнуляет аргумент «одна видеокарта лучше двух».
https://habr.com/ru/articles/1088316/
#llamacpp #локальные_llm #moe #gguf #инференс_llm #vram_использование #pcie #qwen36 #Qwen_36 #локальный_ии
-
RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…
mehr auf Arint.info
#AGENT #Agent #agent #Apple #llama #nitter #Qwen36 #qwen36 #Wikipedia #arint_info
-
RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…
mehr auf Arint.info
#AGENT #Agent #agent #Apple #llama #nitter #Qwen36 #qwen36 #Wikipedia #arint_info
-
RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…
mehr auf Arint.info
#AGENT #Agent #agent #Apple #llama #nitter #Qwen36 #qwen36 #Wikipedia #arint_info
-
RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…
mehr auf Arint.info
#Agent #AGENT #agent #Apple #llama #nitter #qwen36 #Qwen36 #Wikipedia #arint_info
-
RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…
mehr auf Arint.info
#Agent #AGENT #agent #Apple #llama #nitter #qwen36 #Qwen36 #Wikipedia #arint_info
-
RT @DataChaz: HOLY SMOKES. A NEW OPEN-SOURCE AGENT JUST BEAT HERMES ON THE GAIA BENCHMARK, RUNNING ON THE EXACT SAME LOCAL MODEL AND HARDWARE @atomicagent_io ran 53 real-world GAIA Level 1 tasks against Hermes using a 4-bit Qwen-3.6-35b on an M4 Max 🤯 The results highlight how much the orchestration layer matters: → Atomic Agent: 69.8% solved (3h 12m) → Hermes Agent: 58.5% solved (5h 10m) Atomic solved 6 more tasks and finished nearly two hours faster. The secret? A highly disciplined agent loop that refuses to waste compute. Atomic uses a byte-stable prompt to massively reuse the KV-cache. Instead of dumping raw logs into the context window, it batches tool calls via JSON and compresses the results. Add in a hard stop for endless tool-call loops, and you get a model that stays razor-sharp instead of drowning in its own junk data. Open-source and local-first! Repo below ↓ Video Atomic Agent (@atomicagent_io) Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a bl…
mehr auf Arint.info
#Agent #AGENT #agent #Apple #llama #nitter #qwen36 #Qwen36 #Wikipedia #arint_info
-
Мини‑ПК на Strix Halo под параллельной нагрузкой: 236 tok/s на 32 одновременных запросах и три ошибки
Один Beelink GTR9 Pro на Ryzen AI Max+ 395 выдал 236 tok/s суммарной генерации на 32 одновременных запросах в коротких прогонах и удержал в среднем 226 tok/s за 30 минут непрерывной нагрузки без тротлинга. По дороге к этим числам я нашёл воспроизводимый провал пропускной способности, который сначала выглядел свойством одной модели, увидел, как спекулятивный декодинг на моём стеке превращается из ускорителя в налог, и трижды чуть не опубликовал неверные выводы. Каждый раз спасали контрольные замеры, все три истории здесь, в статье. Харнессы, конфиги, полные таблицы, значения по каждому ключевому и финальному прогону, поминутные ряды выносливости и телеметрия лежат в открытом репозитории . Логи отдельных запросов харнесс не вёл, поэтому пересчитать можно всё до уровня прогона, но не глубже. Числа сняты на одном конкретном стенде; что из этого переносится на другие стеки, а что нет, оговорено по ходу текста.
https://habr.com/ru/articles/1060520/
#Strix_Halo #Ryzen_AI_Max #llamacpp #локальные_LLM #инференс #бенчмарки #Vulkan #Gemma_4 #Qwen36 #спекулятивный_декодинг
-
Мини‑ПК на Strix Halo под параллельной нагрузкой: 236 tok/s на 32 одновременных запросах и три ошибки
Один Beelink GTR9 Pro на Ryzen AI Max+ 395 выдал 236 tok/s суммарной генерации на 32 одновременных запросах в коротких прогонах и удержал в среднем 226 tok/s за 30 минут непрерывной нагрузки без тротлинга. По дороге к этим числам я нашёл воспроизводимый провал пропускной способности, который сначала выглядел свойством одной модели, увидел, как спекулятивный декодинг на моём стеке превращается из ускорителя в налог, и трижды чуть не опубликовал неверные выводы. Каждый раз спасали контрольные замеры, все три истории здесь, в статье. Харнессы, конфиги, полные таблицы, значения по каждому ключевому и финальному прогону, поминутные ряды выносливости и телеметрия лежат в открытом репозитории . Логи отдельных запросов харнесс не вёл, поэтому пересчитать можно всё до уровня прогона, но не глубже. Числа сняты на одном конкретном стенде; что из этого переносится на другие стеки, а что нет, оговорено по ходу текста.
https://habr.com/ru/articles/1060520/
#Strix_Halo #Ryzen_AI_Max #llamacpp #локальные_LLM #инференс #бенчмарки #Vulkan #Gemma_4 #Qwen36 #спекулятивный_декодинг
-
ИИ Qwen3.6-27B запустили на смартфоне: 1 бит на вес и 90% интеллекта оригинала
Стартап PrismML представил Bonsai 27B — сжатые версии открытой модели Qwen3.6-27B, младшая из которых стала первой нейросетью такого класса, которая помещается в память смартфона. Веса выложены на Hugging Face под лицензией Apache 2.0, а в демонстрациях PrismML модель работает прямо на iPhone 17 Pro Max — рассуждает, вызывает инструменты и разбирает скриншоты без единого обращения к облаку.
-
ИИ Qwen3.6-27B запустили на смартфоне: 1 бит на вес и 90% интеллекта оригинала
Стартап PrismML представил Bonsai 27B — сжатые версии открытой модели Qwen3.6-27B, младшая из которых стала первой нейросетью такого класса, которая помещается в память смартфона. Веса выложены на Hugging Face под лицензией Apache 2.0, а в демонстрациях PrismML модель работает прямо на iPhone 17 Pro Max — рассуждает, вызывает инструменты и разбирает скриншоты без единого обращения к облаку.
-
RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x
mehr auf Arint.info
-
RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x
mehr auf Arint.info
-
RT @spiritbuun: My previous best decode on a single 3090 with Qwen3.6 27B was 206 tok/s. Today I beat it. 219 tok/s. One 3090. DFlash perf work has been PUSHED, get it now on buun-llama. Several more DFlash improvements are in development- expect several more incremental improvements this week
mehr auf Arint.info
-
RT @spiritbuun: My previous best decode on a single 3090 with Qwen3.6 27B was 206 tok/s. Today I beat it. 219 tok/s. One 3090. DFlash perf work has been PUSHED, get it now on buun-llama. Several more DFlash improvements are in development- expect several more incremental improvements this week
mehr auf Arint.info
-
RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x
mehr auf Arint.info
-
RT @superalesha: qwen3.6 27b just decoded at 150+ tok/s on two used rtx 3090s not the moe. the DENSE 27b, the 24gb-tier king. no blackwell, no nvlink, no fp8. the trick is dflash: a small draft model guesses 15 tokens ahead and the 27b verifies the whole guess in one pass. easy math answers burst past 160 cause the guesses keep landing. my real coding prompts: 54 -> 90 tok/s median. same cards, same weights, one draft model bolted on. a clean 1.67x
mehr auf Arint.info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
RT @MiaAI_lab: Welches lokale Modell ist das beste für Agentic Workflows auf einem einzelnen @NVIDIAAI DGX Spark? (oder einem anderen Rig mit 96-128 GB VRAM) Nach der Auswertung von 84 Szenarien, 16 Kategorien und jeweils 8 Durchläufen im Hermes-Agent-Stil für die mehrstufige Tool-Orchestrierung gibt es einen sehr klaren Sieger. 🏆 Qwen 3.6 35B A3B Q8KXL liegt auf Platz 1. Es ist das einzige Modell, das durchweg perfekte Ergebnisse erzielte und keine katastrophalen Ausfälle aufwies. Die vollständige Rangliste: Qwen 3.6 35B A3B UD Q8KXL — 91,0 | Qwen 3.6 27B NVFP4 — 89,0 | Qwopus 3.6 27B Coder MTP — 85,2 | DeepSeek V4 Flash Q2 — 86,5 | Agents-A1 Q80 — 83,4 | Gemma 4 26B — 81,4 | Nemotron 3 Nano Omni 30B — 79,0 Fazit: Wenn du 2026 Agenten lokal auf einem DGX Spark oder einem anderen Rig mit 96–128 GB VRAM betreibst, ist Qwen 3.6 35B Q8KXL derzeit die beste Wahl. Vollständiger Bericht + tiefgehende Analyse 👇 https://github.com/MiaAI-Lab/Best-Local-ModelAgentic-Workflows2026
mehr auf Arint.info
#AgenticAI #AIAgents #LocalLLM #MachineLearning #NVIDIADGX #Qwen36 #arint_info
-
RT @MiaAI_lab: Welches lokale Modell ist das beste für Agentic Workflows auf einem einzelnen @NVIDIAAI DGX Spark? (oder einem anderen Rig mit 96-128 GB VRAM) Nach der Auswertung von 84 Szenarien, 16 Kategorien und jeweils 8 Durchläufen im Hermes-Agent-Stil für die mehrstufige Tool-Orchestrierung gibt es einen sehr klaren Sieger. 🏆 Qwen 3.6 35B A3B Q8KXL liegt auf Platz 1. Es ist das einzige Modell, das durchweg perfekte Ergebnisse erzielte und keine katastrophalen Ausfälle aufwies. Die vollständige Rangliste: Qwen 3.6 35B A3B UD Q8KXL — 91,0 | Qwen 3.6 27B NVFP4 — 89,0 | Qwopus 3.6 27B Coder MTP — 85,2 | DeepSeek V4 Flash Q2 — 86,5 | Agents-A1 Q80 — 83,4 | Gemma 4 26B — 81,4 | Nemotron 3 Nano Omni 30B — 79,0 Fazit: Wenn du 2026 Agenten lokal auf einem DGX Spark oder einem anderen Rig mit 96–128 GB VRAM betreibst, ist Qwen 3.6 35B Q8KXL derzeit die beste Wahl. Vollständiger Bericht + tiefgehende Analyse 👇 https://github.com/MiaAI-Lab/Best-Local-ModelAgentic-Workflows2026
mehr auf Arint.info
#AgenticAI #AIAgents #LocalLLM #MachineLearning #NVIDIADGX #Qwen36 #arint_info
-
RT @Tono_Ken3: Und bei StrixHalo's DearfStar4 läuft DeepSeek-V4-Flash mit 16 TPS. Da der Prefill den KV-Cache auf einer Optane-SSD speichert, ist die Geschwindigkeit wirklich beeindruckend. Es könnte auch gut sein, diesen Hermes als Sub-Agent von Qwen3.6's Lnagent aufzurufen. Das lokale Agenten-System besteht aus diesen drei Modellen: Qwen3.6-35b-a3b-nvfp4, DeepSeek-V4Flash-IQ2 und GLM-5.2-UQ4. Es ist übersichtlich. TonoKen3🤖Local-LLM&Robot🏁とのけん3 (@TonoKen3) Ja genau. Die Möglichkeit, GLM-5.2 lokal einzusetzen, schafft ein Gefühl von innerem Frieden. Für 90% der Fälle reicht die schnelle Antwort von Qwen3.6-35b. 130 TPS bieten eine komfortable Reaktionsgeschwindigkeit, die sogar die Nutzung geschlossener Modelle übertrifft. Bei der Inferenz verbraucht das System 550W, im Standby nur 200W. Das ist genau das, wonach man sucht. — https://nitter.net/TonoKen3/status/2073892875117691049#m
mehr auf Arint.info
#AIInfrastructure #DeepSeekV4 #GLM52 #LocalLLM #Qwen36 #TonoKen3 #arint_info
-
RT @Tono_Ken3: Und bei StrixHalo's DearfStar4 läuft DeepSeek-V4-Flash mit 16 TPS. Da der Prefill den KV-Cache auf einer Optane-SSD speichert, ist die Geschwindigkeit wirklich beeindruckend. Es könnte auch gut sein, diesen Hermes als Sub-Agent von Qwen3.6's Lnagent aufzurufen. Das lokale Agenten-System besteht aus diesen drei Modellen: Qwen3.6-35b-a3b-nvfp4, DeepSeek-V4Flash-IQ2 und GLM-5.2-UQ4. Es ist übersichtlich. TonoKen3🤖Local-LLM&Robot🏁とのけん3 (@TonoKen3) Ja genau. Die Möglichkeit, GLM-5.2 lokal einzusetzen, schafft ein Gefühl von innerem Frieden. Für 90% der Fälle reicht die schnelle Antwort von Qwen3.6-35b. 130 TPS bieten eine komfortable Reaktionsgeschwindigkeit, die sogar die Nutzung geschlossener Modelle übertrifft. Bei der Inferenz verbraucht das System 550W, im Standby nur 200W. Das ist genau das, wonach man sucht. — https://nitter.net/TonoKen3/status/2073892875117691049#m
mehr auf Arint.info
#AIInfrastructure #DeepSeekV4 #GLM52 #LocalLLM #Qwen36 #TonoKen3 #arint_info
-
Контекстная инженерия для слабой локальной модели: как мы делаем среднюю модель надёжной
Принято думать, что качество ИИ-агента упирается в размер модели. Но когда модель работает локально, в закрытом контуре и на ограниченном железе, брать «побольше» особо некуда. И оказывается, что главный рычаг не модель, а контекст: что вы ей показываете, в каком порядке и как фильтруете. Причём «контекст» здесь — это сборка под то, кто спрашивает, откуда и о чём, плюс честный порог релевантности и продуманный порядок секций. На сильной облачной модели небрежный контекст прощается запасом по reasoning; на средней локальной — нет. Об этом и статья.
https://habr.com/ru/companies/1forma/articles/1054228/
#llm #onpremise #qwen36 #ai #aiагенты #искусственный_интеллект #автоматизация_процессов #корпоративные_системы #lowcode #bpms
-
Контекстная инженерия для слабой локальной модели: как мы делаем среднюю модель надёжной
Принято думать, что качество ИИ-агента упирается в размер модели. Но когда модель работает локально, в закрытом контуре и на ограниченном железе, брать «побольше» особо некуда. И оказывается, что главный рычаг не модель, а контекст: что вы ей показываете, в каком порядке и как фильтруете. Причём «контекст» здесь — это сборка под то, кто спрашивает, откуда и о чём, плюс честный порог релевантности и продуманный порядок секций. На сильной облачной модели небрежный контекст прощается запасом по reasoning; на средней локальной — нет. Об этом и статья.
https://habr.com/ru/companies/1forma/articles/1054228/
#llm #onpremise #qwen36 #ai #aiагенты #искусственный_интеллект #автоматизация_процессов #корпоративные_системы #lowcode #bpms
-
So I saw this Spider-man reference over on Reddit, and I realized I wasn't familiar with what Peter is referencing here.
https://en.wikipedia.org/wiki/Brachistochrone_curve
It's pretty neat.
But I also had a side thought: let's throw it at #Qwen36 and have it make an interactive demonstration?
After about ~15 mins of churn, it made a single file HTML: https://scratch.network47.org/s/a8yocvi7vf
This is WITHOUT using Wikipedia as a reference. Purely from the model.
-
So I saw this Spider-man reference over on Reddit, and I realized I wasn't familiar with what Peter is referencing here.
https://en.wikipedia.org/wiki/Brachistochrone_curve
It's pretty neat.
But I also had a side thought: let's throw it at #Qwen36 and have it make an interactive demonstration?
After about ~15 mins of churn, it made a single file HTML: https://scratch.network47.org/s/a8yocvi7vf
This is WITHOUT using Wikipedia as a reference. Purely from the model.
-
So I saw this Spider-man reference over on Reddit, and I realized I wasn't familiar with what Peter is referencing here.
https://en.wikipedia.org/wiki/Brachistochrone_curve
It's pretty neat.
But I also had a side thought: let's throw it at #Qwen36 and have it make an interactive demonstration?
After about ~15 mins of churn, it made a single file HTML: https://scratch.network47.org/s/a8yocvi7vf
This is WITHOUT using Wikipedia as a reference. Purely from the model.
-
So I saw this Spider-man reference over on Reddit, and I realized I wasn't familiar with what Peter is referencing here.
https://en.wikipedia.org/wiki/Brachistochrone_curve
It's pretty neat.
But I also had a side thought: let's throw it at #Qwen36 and have it make an interactive demonstration?
After about ~15 mins of churn, it made a single file HTML: https://scratch.network47.org/s/a8yocvi7vf
This is WITHOUT using Wikipedia as a reference. Purely from the model.
-
Qwen 3.6 27B is the sweet spot for local development
https://quesma.com/blog/qwen-36-is-awesome/
#HackerNews #Qwen36 #LocalDevelopment #SoftwareDevelopment #TechNews #CodingInsights
-
Qwen 3.6 27B is the sweet spot for local development
https://quesma.com/blog/qwen-36-is-awesome/
#HackerNews #Qwen36 #LocalDevelopment #SoftwareDevelopment #TechNews #CodingInsights
-
Qwen 3.6 27B is the sweet spot for local development
https://quesma.com/blog/qwen-36-is-awesome/
#HackerNews #Qwen36 #LocalDevelopment #SoftwareDevelopment #TechNews #CodingInsights
-
Qwen 3.6 27B is the sweet spot for local development
https://quesma.com/blog/qwen-36-is-awesome/
#HackerNews #Qwen36 #LocalDevelopment #SoftwareDevelopment #TechNews #CodingInsights
-
Тесты бюджетных сборок для ИИ до 100к рублей
Локальный ИИ не должен стоить как автомобиль. Мне стало интересно: возможен ли жизнеспособный инференс на CPU и что реально дают дешевые GPU (вроде Tesla V100 или CMP 40HX). Я собрал несколько бюджетных конфигураций до 100к, потестил актуальные модели и попытался понять, что важнее для скорости: канальность памяти или частота. Сравнил дешевые AM4 и Threadripper, замерил токены в секунду и построил графики. Делюсь результатами.
https://habr.com/ru/articles/1053118/
#ai #ии #gpt #selfhosted #gpu #cpu #llamacpp #qwen36 #gemma4
-
Тесты бюджетных сборок для ИИ до 100к рублей
Локальный ИИ не должен стоить как автомобиль. Мне стало интересно: возможен ли жизнеспособный инференс на CPU и что реально дают дешевые GPU (вроде Tesla V100 или CMP 40HX). Я собрал несколько бюджетных конфигураций до 100к, потестил актуальные модели и попытался понять, что важнее для скорости: канальность памяти или частота. Сравнил дешевые AM4 и Threadripper, замерил токены в секунду и построил графики. Делюсь результатами.
https://habr.com/ru/articles/1053118/
#ai #ии #gpt #selfhosted #gpu #cpu #llamacpp #qwen36 #gemma4
-
RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
https://imil.net/blog/posts/2026/rtx-5080-+-rtx-3090-setup-80+-tok-s-on-qwen-3.6-27b-q8/
#HackerNews #RTX5080 #RTX3090 #Qwen36 #TokPerformance #AIComputing
-
RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
https://imil.net/blog/posts/2026/rtx-5080-+-rtx-3090-setup-80+-tok-s-on-qwen-3.6-27b-q8/
#HackerNews #RTX5080 #RTX3090 #Qwen36 #TokPerformance #AIComputing
-
RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
https://imil.net/blog/posts/2026/rtx-5080-+-rtx-3090-setup-80+-tok-s-on-qwen-3.6-27b-q8/
#HackerNews #RTX5080 #RTX3090 #Qwen36 #TokPerformance #AIComputing
-
RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
https://imil.net/blog/posts/2026/rtx-5080-+-rtx-3090-setup-80+-tok-s-on-qwen-3.6-27b-q8/
#HackerNews #RTX5080 #RTX3090 #Qwen36 #TokPerformance #AIComputing
-
Локальные LLM на Arch Linux и как увеличить скорость генерации в 20 раз
Приветствую всех читателей Хабра, в этой статье я хочу поделиться своим опытом в запуске локальных LLM, протестировать работоспособность интересных моделей на своем железе, рассказать, как я увеличил скорость генерации на одной из нейросетей в 20 раз (я не преувеличиваю). Но об этом чуть позже, а начну я повествование с описания своего железа.
https://habr.com/ru/articles/1045898/
#arch_linux #llamacpp #ollama #qwen36 #gemma4 #github #huggingface #intel_arc_b580
-
Локальные LLM на Arch Linux и как увеличить скорость генерации в 20 раз
Приветствую всех читателей Хабра, в этой статье я хочу поделиться своим опытом в запуске локальных LLM, протестировать работоспособность интересных моделей на своем железе, рассказать, как я увеличил скорость генерации на одной из нейросетей в 20 раз (я не преувеличиваю). Но об этом чуть позже, а начну я повествование с описания своего железа.
https://habr.com/ru/articles/1045898/
#arch_linux #llamacpp #ollama #qwen36 #gemma4 #github #huggingface #intel_arc_b580
-
Claude Code с локальными Qwen3.6 на AMD Strix Halo: полное руководство по настройке
Всем привет! Продолжаю тему локальных LLM. В предыдущей статье мы сравнивали железо для инференса — Nvidia DGX Spark, Mac Studio M3 Ultra и Strix Halo. И как можно было догадаться, я остановился именно на последнем. Теперь, когда железка есть, встает вопрос: а как из неё извлечь практическую пользу? Claude code с оригинальными LLM - это, конечно, замечательно. Но это стоят денег, да и свой код в чужие дата-центры не всегда правильно лить. Плюс за всякое неосторожное движение можно попасть в бан, рискуя потерять все свои наработки. Одно из решений: Claude Code во free mode с локальными моделями . Anthropic позволяет заменить свои модели на любые с совместимым API. То есть, на что угодно — даже на модель, крутящуюся прямо у вас на компьютере. В этой статье я расскажу, как всё это настроить на Strix Halo — от загрузки моделей до первого запроса к Claude Code.
https://habr.com/ru/articles/1038774/
#claudecode #strix_halo #ииагенты #программирование #antropic #qwen36 #локальный_ии #llamacpp #vibecoding
-
Claude Code с локальными Qwen3.6 на AMD Strix Halo: полное руководство по настройке
Всем привет! Продолжаю тему локальных LLM. В предыдущей статье мы сравнивали железо для инференса — Nvidia DGX Spark, Mac Studio M3 Ultra и Strix Halo. И как можно было догадаться, я остановился именно на последнем. Теперь, когда железка есть, встает вопрос: а как из неё извлечь практическую пользу? Claude code с оригинальными LLM - это, конечно, замечательно. Но это стоят денег, да и свой код в чужие дата-центры не всегда правильно лить. Плюс за всякое неосторожное движение можно попасть в бан, рискуя потерять все свои наработки. Одно из решений: Claude Code во free mode с локальными моделями . Anthropic позволяет заменить свои модели на любые с совместимым API. То есть, на что угодно — даже на модель, крутящуюся прямо у вас на компьютере. В этой статье я расскажу, как всё это настроить на Strix Halo — от загрузки моделей до первого запроса к Claude Code.
https://habr.com/ru/articles/1038774/
#claudecode #strix_halo #ииагенты #программирование #antropic #qwen36 #локальный_ии #llamacpp #vibecoding
-
Tesla v100 SXM2 X2 32GB total
В этом материале я разбираю практический кейс: развёртывание Qwen3.6-27B на двух Tesla V100-SXM2-16GB под управлением автономного агента Hermes от Nous Research. Карты подключены к потребительской платформе через адаптеры SXM2→PCIe — конфигурация, которую несложно собрать дома, но которая накладывает жёсткие ограничения на доступную видеопамять и межкарточную пропускную способность.
https://habr.com/ru/articles/1043956/
#tesla_v100 #v100 #SXM2 #qwen #qwen36 #2017
-
Tesla v100 SXM2 X2 32GB total
В этом материале я разбираю практический кейс: развёртывание Qwen3.6-27B на двух Tesla V100-SXM2-16GB под управлением автономного агента Hermes от Nous Research. Карты подключены к потребительской платформе через адаптеры SXM2→PCIe — конфигурация, которую несложно собрать дома, но которая накладывает жёсткие ограничения на доступную видеопамять и межкарточную пропускную способность.
https://habr.com/ru/articles/1043956/
#tesla_v100 #v100 #SXM2 #qwen #qwen36 #2017
-
RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…
mehr auf Arint.info
#Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info
-
RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…
mehr auf Arint.info
#Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info
-
RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…
mehr auf Arint.info
#Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info
-
RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…
mehr auf Arint.info
#Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info
-
RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…
mehr auf Arint.info
#Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info
-
Почему Qwen3.6-27B лучше чем Claude? Железная коробка, которая научилась думать
На вопрос «Чем локальная модель лучше коммерческой top‑quality модели от Anthropic, OpenAI или Google?», — обычно отвечают: приватность. На самом деле это не совсем так. Приватность важна, но не только она. У локальных моделей есть более важные качества, которые я покажу в этой статье.
-
Почему Qwen3.6-27B лучше чем Claude? Железная коробка, которая научилась думать
На вопрос «Чем локальная модель лучше коммерческой top‑quality модели от Anthropic, OpenAI или Google?», — обычно отвечают: приватность. На самом деле это не совсем так. Приватность важна, но не только она. У локальных моделей есть более важные качества, которые я покажу в этой статье.
-
RT @AtlasInference: TRANSLASATION: DGX Spark hat gerade für Qwen3.6-35B mit @AtlasInference auf @sparkarena über 200 Token pro Sekunde erreicht 🔥
mehr auf Arint.info
#AIInnovation #AtlasInference #DGXSpark #LLMPerformance #Qwen36 #TokenSpeed #arint_info
-
Qwen3.6 MTP весит на 0.3 Гб больше, а даёт ускорение в ~2 раза. С 60 t/s до 130 t/s для Qwen3.6 27B без искажений
В llama.cpp добавили поддержку MTP Qwen3.6. Дополнительные слои Multi-Token Prediction позволяют сгенерировать сразу несколько токенов за 1 проход, что ускоряет генерацию в 1.5-2 раза. Качество при этом остается lossless. Для моделей, которые не имеют встроенного MTP, есть альтернативы в лице EAGLE-3 и DFlash.
-
Qwen3.6 MTP весит на 0.3 Гб больше, а даёт ускорение в ~2 раза. С 60 t/s до 130 t/s для Qwen3.6 27B без искажений
В llama.cpp добавили поддержку MTP Qwen3.6. Дополнительные слои Multi-Token Prediction позволяют сгенерировать сразу несколько токенов за 1 проход, что ускоряет генерацию в 1.5-2 раза. Качество при этом остается lossless. Для моделей, которые не имеют встроенного MTP, есть альтернативы в лице EAGLE-3 и DFlash.