home.social

#qwen35 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #qwen35, aggregated by home.social.

fetched live
  1. Локальный запуск LLM для SOC: сколько инцидентов обработает одна GPU? Часть 2

    Всем привет! На связи Сергей Иванов, аналитик технологий машинного обучения R‑Vision. В первой части эксперимента мы выяснили, как на производительность локальной LLM влияют длина контекста, количество параллельных запросов и объем генерируемого ответа. Стресс‑тесты позволили определить границы конфигурации Qwen3.5–122B‑A10B‑GPTQ, vLLM и NVIDIA RTX PRO 6000 Blackwell Max‑Q с 96 GB видеопамяти. Однако предельная конкурентность и скорость генерации сами по себе еще не показывают, насколько такая конфигурация подходит для реального SOC. В промышленном сценарии запросы поступают не равномерно и не изолированно. Они создаются карточками инцидентов, шагами ИИ‑оркестратора и действиями аналитиков, а порядок их выполнения определяется логикой расследования. Во второй части эксперимента мы перешли от лабораторных измерений к моделированию реальной работы SOC. Мы оценили, как GPU справляется с инференсом LLM при разной численности команды и интенсивности потока инцидентов — от спокойной смены до пиковых ситуаций с массовым поступлением новых инцидентов. Важно отметить, что в эксперименте использовались не специально подготовленные тестовые примеры, а анонимизированные реальные инциденты из практики нашего внутреннего SOC. Отдельно остановимся на режиме рассуждений. В сценарии интеграции LLM в конвейер R‑Vision SOAR мы сознательно использовали модель с отключенным thinking mode (режимом рассуждений). Ниже на результатах реальных экспериментов покажем, почему именно такой режим оказался наиболее эффективным для задач SOC.

    habr.com/ru/companies/rvision/

    #llm #soc #gpu #автоматизация_SOC #nvidia_rtx_pro_6000_blackwell #vllm #qwen35 #AI_в_SOC #инференс_llm #selfhosted_llm

  2. Локальный запуск LLM для SOC: сколько инцидентов обработает одна GPU? Часть 2

    Всем привет! На связи Сергей Иванов, аналитик технологий машинного обучения R‑Vision. В первой части эксперимента мы выяснили, как на производительность локальной LLM влияют длина контекста, количество параллельных запросов и объем генерируемого ответа. Стресс‑тесты позволили определить границы конфигурации Qwen3.5–122B‑A10B‑GPTQ, vLLM и NVIDIA RTX PRO 6000 Blackwell Max‑Q с 96 GB видеопамяти. Однако предельная конкурентность и скорость генерации сами по себе еще не показывают, насколько такая конфигурация подходит для реального SOC. В промышленном сценарии запросы поступают не равномерно и не изолированно. Они создаются карточками инцидентов, шагами ИИ‑оркестратора и действиями аналитиков, а порядок их выполнения определяется логикой расследования. Во второй части эксперимента мы перешли от лабораторных измерений к моделированию реальной работы SOC. Мы оценили, как GPU справляется с инференсом LLM при разной численности команды и интенсивности потока инцидентов — от спокойной смены до пиковых ситуаций с массовым поступлением новых инцидентов. Важно отметить, что в эксперименте использовались не специально подготовленные тестовые примеры, а анонимизированные реальные инциденты из практики нашего внутреннего SOC. Отдельно остановимся на режиме рассуждений. В сценарии интеграции LLM в конвейер R‑Vision SOAR мы сознательно использовали модель с отключенным thinking mode (режимом рассуждений). Ниже на результатах реальных экспериментов покажем, почему именно такой режим оказался наиболее эффективным для задач SOC.

    habr.com/ru/companies/rvision/

    #llm #soc #gpu #автоматизация_SOC #nvidia_rtx_pro_6000_blackwell #vllm #qwen35 #AI_в_SOC #инференс_llm #selfhosted_llm

  3. RT @TheAhmadOsman: Ich habe gerade mit @PrismMLs neuem Modell experimentiert, das Qwen 3.5 27B in Unter-4GB- und Unter-6GB-Gewichte umgewandelt hat, und ich bin beeindruckt. Ich kann kaum glauben, wie weit Open-Source und lokale KI seit Weihnachten (vor etwa 8 Monaten) gekommen sind.

    mehr auf Arint.info

    #AIResearch #LocalAI #MachineLearning #OpenSource #PrismML #Qwen35 #arint_info

    https://x.com/TheAhmadOsman/status/2077536563303457028#m

  4. RT @TheAhmadOsman: Ich habe gerade mit @PrismMLs neuem Modell experimentiert, das Qwen 3.5 27B in Unter-4GB- und Unter-6GB-Gewichte umgewandelt hat, und ich bin beeindruckt. Ich kann kaum glauben, wie weit Open-Source und lokale KI seit Weihnachten (vor etwa 8 Monaten) gekommen sind.

    mehr auf Arint.info

    #AIResearch #LocalAI #MachineLearning #OpenSource #PrismML #Qwen35 #arint_info

    https://x.com/TheAhmadOsman/status/2077536563303457028#m

  5. RT @songqiaosu: 🐦 Ornith-1.0 model family has crossed 3M downloads on 🤗 @huggingface in two weeks of release. This milestone belongs to the community! Please leave any feedback in the comments! Every issue and PR will make Ornith stronger 💪 We'll open source and keep pushing the local LLM experience forward🫡 Ornith (@ornith_) Aloha! 🌺 Meet Ornith-1.0, a family of open-source LLMs specialized for agentic coding. Ornith-1.0 spans the full parameter sizes including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. It achieves state-of-the-art performance among open-source models of comparable size on coding benchmarks including: ✅Terminal-Bench 2.1(77.5) ✅SWE-Bench(82.4 on verified, 62.2 on pro, 78.9 on Multilingual) ✅NL2Repo(48.2) ✅SWE Atlas(41.2 on QnA, 42.6 RF, 39.1 TW) ✅ClawEval(77.1) Post-trained on top of gemma4 and qwen3.5, Ornith-1.0 employs a novel self-improving training strategy in which reinforcement learning is used to generate not only solution rollouts, but also the task-specific scaffolds that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model generate higher-quality solutions in agentic coding.😎 All models are released under the MIT license, enabling full commercial and research use. 📖Tech Blog: deep-reinforce.com/ornith_1_… 🤗Huggingface: huggingface.co/collections/d… — nitter.net/ornith_/status/2070

    mehr auf Arint.info

    #huggingface #Huggingface #make #MIT #MITlicense #nitter #opensource #qwen35 #SWE #SWEBench #arint_info

    https://x.com/songqiaosu/status/2076743265328726034#m

  6. RT @songqiaosu: 🐦 Ornith-1.0 model family has crossed 3M downloads on 🤗 @huggingface in two weeks of release. This milestone belongs to the community! Please leave any feedback in the comments! Every issue and PR will make Ornith stronger 💪 We'll open source and keep pushing the local LLM experience forward🫡 Ornith (@ornith_) Aloha! 🌺 Meet Ornith-1.0, a family of open-source LLMs specialized for agentic coding. Ornith-1.0 spans the full parameter sizes including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. It achieves state-of-the-art performance among open-source models of comparable size on coding benchmarks including: ✅Terminal-Bench 2.1(77.5) ✅SWE-Bench(82.4 on verified, 62.2 on pro, 78.9 on Multilingual) ✅NL2Repo(48.2) ✅SWE Atlas(41.2 on QnA, 42.6 RF, 39.1 TW) ✅ClawEval(77.1) Post-trained on top of gemma4 and qwen3.5, Ornith-1.0 employs a novel self-improving training strategy in which reinforcement learning is used to generate not only solution rollouts, but also the task-specific scaffolds that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model generate higher-quality solutions in agentic coding.😎 All models are released under the MIT license, enabling full commercial and research use. 📖Tech Blog: deep-reinforce.com/ornith_1_… 🤗Huggingface: huggingface.co/collections/d… — nitter.net/ornith_/status/2070

    mehr auf Arint.info

    #huggingface #Huggingface #make #MIT #MITlicense #nitter #opensource #qwen35 #SWE #SWEBench #arint_info

    https://x.com/songqiaosu/status/2076743265328726034#m

  7. Vera возвращается: как голосовой ассистент превратился в локального AI-агента для Windows

    Я собирался просто добавить Vera графический интерфейс, а в итоге переписал почти весь проект. Теперь она умеет работать с файлами и изображениями, запоминать пользователя, выполнять фоновые задачи, создавать презентации и все это с одной небольшой локальной моделью.

    habr.com/ru/articles/1058810/

    #агент #qwen35 #искусственный_интеллект #пк #python #vera

  8. Vera возвращается: как голосовой ассистент превратился в локального AI-агента для Windows

    Я собирался просто добавить Vera графический интерфейс, а в итоге переписал почти весь проект. Теперь она умеет работать с файлами и изображениями, запоминать пользователя, выполнять фоновые задачи, создавать презентации и все это с одной небольшой локальной моделью.

    habr.com/ru/articles/1058810/

    #агент #qwen35 #искусственный_интеллект #пк #python #vera

  9. Qwen3.5 на двух V100, reverse SSH вместо Cloudflare в Telegram Mini App: собираю AI-репетитора английского

    У меня в углу комнаты стоит сервер с двумя Tesla V100 32GB. Они доcтались мне для другой задачи, которая отвалилась, и полгода стояли мёртвым грузом. Параллельно я в очередной раз пробовал заниматься английским — Simpler, Doalingo, ещё пара продуктов. Хорошие, но мне не подходил формат: я хотел сценарий «открыл телефон дома на семь минут, поговорил, закрыл». Без расписания, без камеры, без поиска тьютора, который понимает мой акцент с пятого раза. Сошлось. Идея: Telegram Mini App, в нём кнопка «говорить», за ней — AI-репетитор, который слышит, что я сказал, отвечает голосом, помнит контекст разговора, тыкает в мои повторяющиеся ошибки и подбрасывает слова, которые я пытаюсь выучить. Полностью бесплатно. Модель Qwen3.5 вышла 25 февраля , я её гоняю всего несколько недель, продукт сырой. Эта статья — про архитектурные решения и про то, на какие грабли я уже успел наступить.

    habr.com/ru/articles/1042166/

    #vllm #qwen35 #telegram_bot #telegram_mini_apps #aiogram_3 #fastapi #selfhosted_llm #kokoro_tts #whisper #tesla_v100

  10. Qwen3.5 на двух V100, reverse SSH вместо Cloudflare в Telegram Mini App: собираю AI-репетитора английского

    У меня в углу комнаты стоит сервер с двумя Tesla V100 32GB. Они доcтались мне для другой задачи, которая отвалилась, и полгода стояли мёртвым грузом. Параллельно я в очередной раз пробовал заниматься английским — Simpler, Doalingo, ещё пара продуктов. Хорошие, но мне не подходил формат: я хотел сценарий «открыл телефон дома на семь минут, поговорил, закрыл». Без расписания, без камеры, без поиска тьютора, который понимает мой акцент с пятого раза. Сошлось. Идея: Telegram Mini App, в нём кнопка «говорить», за ней — AI-репетитор, который слышит, что я сказал, отвечает голосом, помнит контекст разговора, тыкает в мои повторяющиеся ошибки и подбрасывает слова, которые я пытаюсь выучить. Полностью бесплатно. Модель Qwen3.5 вышла 25 февраля , я её гоняю всего несколько недель, продукт сырой. Эта статья — про архитектурные решения и про то, на какие грабли я уже успел наступить.

    habr.com/ru/articles/1042166/

    #vllm #qwen35 #telegram_bot #telegram_mini_apps #aiogram_3 #fastapi #selfhosted_llm #kokoro_tts #whisper #tesla_v100

  11. RT @ArtificialAnlys: OpenBMB has released MiniCPM5-1B (Non-reasoning), the leading 1B open weights model, scoring 17.9 on the Artificial Analysis Intelligence Index @OpenBMB is a China-based lab jointly founded in 2022 by Tsinghua University’s NLP Lab and ModelBest Inc. This release extends the open weights Pareto frontier for Intelligence vs. Parameters at the sub-2B scale. It sits almost 2 points ahead of the best-performing 2B open weights model, @Alibaba's Qwen3.5 2B (Reasoning, 16.3), and 7 points ahead of Qwen3.5 0.8B (Reasoning, 10.5). Unlike the recently released MiniCPM-V 4.6 1.3B Instruct, MiniCPM5-1B (Non-reasoning) does not support native multimodal input, and is text input and output only. Key results: ➤ MiniCPM5-1B scores 17.9 on the Artificial Analysis Intelligence Index, the highest of any open weights model at 1B parameters or below by 7.4 points. The next-most-intelligent open weights model at this scale is Qwen3.5 0.8B (Reasoning, 10.5). No other open weights model under 2B parameters has exceeded 15 on the Intelligence Index; its predecessor MiniCPM-V 4.6 1.3B sits at 12.7. ➤ MiniCPM5-1B extends the open weights Pareto frontier on both Intelligence vs. Total Parameters and Intelligence vs. Active Parameters at the sub-2B scale. It surpasses its predecessor MiniCPM-V 4.6 1.3B (12.7) by 5.3 points at ~23% fewer parameters, and beats Qwen3.5 2B (Reasoning, 16.3) by 1.6 points at less than half the parameter count. ➤ MiniCPM5-1B is more token-efficient than the larger reasoning peers it surpasses, but uses more output tokens than its (also non-reasoning) predecessor MiniCP…

    mehr auf Arint.info

    #Alibaba #Apache #China #Qwen35 #scale #arint_info

    https://x.com/ArtificialAnlys/status/2059411573907808487#m

  12. RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…

    mehr auf Arint.info

    #Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info

    https://x.com/ChujieZheng/status/2057403166589956518#m

  13. RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…

    mehr auf Arint.info

    #Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info

    https://x.com/ChujieZheng/status/2057403166589956518#m

  14. RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…

    mehr auf Arint.info

    #Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info

    https://x.com/ChujieZheng/status/2057403166589956518#m

  15. RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…

    mehr auf Arint.info

    #Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info

    https://x.com/ChujieZheng/status/2057403166589956518#m

  16. Mechanistic Anatomy of Political Constraint in Qwen 3.5

    New technical research on May 20, 2026, shows Qwen 3.5 has censorship rules built into its core code. Learn how this affects how the AI answers questions.

    #qwen35, #aiethics, #techresearch, #aigovernance, #digitalprivacy

    newsletter.tf/qwen-3-5-ai-cens

  17. Mechanistic Anatomy of Political Constraint in Qwen 3.5

    New technical research on May 20, 2026, shows Qwen 3.5 has censorship rules built into its core code. Learn how this affects how the AI answers questions.

    #qwen35, #aiethics, #techresearch, #aigovernance, #digitalprivacy

    newsletter.tf/qwen-3-5-ai-cens

  18. Mechanistic Anatomy of Political Constraint in Qwen 3.5

    New technical research on May 20, 2026, shows Qwen 3.5 has censorship rules built into its core code. Learn how this affects how the AI answers questions.

    #qwen35, #aiethics, #techresearch, #aigovernance, #digitalprivacy

    newsletter.tf/qwen-3-5-ai-cens

  19. A new study shows the Qwen 3.5 AI model has political rules built directly into its brain. This is different from other AI models that use simple safety filters.

    #qwen35, #aiethics, #techresearch, #aigovernance, #digitalprivacy
    newsletter.tf/qwen-3-5-ai-cens

  20. A new study shows the Qwen 3.5 AI model has political rules built directly into its brain. This is different from other AI models that use simple safety filters.

    #qwen35, #aiethics, #techresearch, #aigovernance, #digitalprivacy
    newsletter.tf/qwen-3-5-ai-cens

  21. A new study shows the Qwen 3.5 AI model has political rules built directly into its brain. This is different from other AI models that use simple safety filters.

    #qwen35, #aiethics, #techresearch, #aigovernance, #digitalprivacy
    newsletter.tf/qwen-3-5-ai-cens

  22. RT @ChujieZheng: For Qwen3.7-Max, we have invested far more compute into RL training than ever before. Its top-tier AA score confirms the resulting general and agentic capabilities. This is just the start. We will firmly push forward RL scaling to build more powerful Qwen models. Stay tuned! Artificial Analysis (@ArtificialAnlys) Alibaba’s new Qwen3.7 Max model scores 56.6 on the Artificial Analysis Intelligence Index, 4.8 points higher than Qwen3.6 Max Preview (51.8). While Alibaba still trails models from OpenAI, Anthropic and Google, Qwen3.7 Max is the closest they have been to the frontier Qwen3.7 Max is @Alibaba_Qwen's latest proprietary flagship, scoring 56.6 on the Intelligence Index, a 4.8 point gain over Qwen3.6 Max Preview (51.8) released in April. Qwen3.7 Max continues Alibaba's pattern, in place since Qwen2.5 Max (January 2025), of releasing Max and Plus models as closed weights while the rest of the Qwen line remains open weights. The leading open weights Qwen on the Intelligence Index is Qwen3.6 27B (Reasoning, 45.8) released in April 2026, and the leading open weights MoE Qwen is Qwen3.5 397B A17B (Reasoning, 45.0) released in February 2026 Key takeaways for the reasoning variant: ➤ The Intelligence Index gains over Qwen3.6 Max Preview are concentrated in scientific reasoning, agentic capability and coding. CritPt +9.7 p.p (3.7% to 13.4%), HLE +9.2 p.p (28.9% to 38.1%), TerminalBench Hard +6.9 p.p (43.9% to 50.8%) and GDPval-AA +42 Elo (1504 to 1546). Scores on other benchmarks in the Intelligence Index are flat compared to Qwen3.6 Max Preview ➤ A significant share of the Int…

    mehr auf Arint.info

    #Alibaba #Anthropic #API #Claude #DeepSeek #Gemini #Google #GPT5 #nitter #OpenAI #Qwen #Qwen25 #Qwen35 #Qwen36 #Qwen37 #rest #arint_info

    https://x.com/ChujieZheng/status/2057403166589956518#m

  23. Топ локальных нейросетей ︎◍ 2026: подборка ИИ для запуска из дома

    Сознаюсь: когда я впервые попытался запустить большую языковую модель на своём ноутбуке, всё закончилось вертушкой кулера, жутким лагом и системным сообщением “Недостаточно памяти”. Казалось, что домашний ИИ – удел владельцев космических станций с жидким азотом. Но прошло совсем немного времени, и ситуация изменилась до неузнаваемости. Теперь достаточно обычной RTX 3060 и получаса свободного вечера, чтобы завести себе персонального ассистента, который работает на даче без интернета и умеет шутить (или хотя бы пытается). Я расскажу обо всём по порядку – без воды и фанатизма. Что вообще запускать, на чём запускать, какие подводные камни ждут и почему “самая новая модель” дома – далеко не всегда лучший выбор. Поехали! Готовьте отвёртку и VRAM – мы начинаем!

    habr.com/ru/companies/bothub/a

    #gemma_4 #qwen36 #qwen35 #gptoss30b #mistral_7b #phi4 #deepseek_v32 #whisper #nemotron_cascade_2

  24. Топ локальных нейросетей ︎◍ 2026: подборка ИИ для запуска из дома

    Сознаюсь: когда я впервые попытался запустить большую языковую модель на своём ноутбуке, всё закончилось вертушкой кулера, жутким лагом и системным сообщением “Недостаточно памяти”. Казалось, что домашний ИИ – удел владельцев космических станций с жидким азотом. Но прошло совсем немного времени, и ситуация изменилась до неузнаваемости. Теперь достаточно обычной RTX 3060 и получаса свободного вечера, чтобы завести себе персонального ассистента, который работает на даче без интернета и умеет шутить (или хотя бы пытается). Я расскажу обо всём по порядку – без воды и фанатизма. Что вообще запускать, на чём запускать, какие подводные камни ждут и почему “самая новая модель” дома – далеко не всегда лучший выбор. Поехали! Готовьте отвёртку и VRAM – мы начинаем!

    habr.com/ru/companies/bothub/a

    #gemma_4 #qwen36 #qwen35 #gptoss30b #mistral_7b #phi4 #deepseek_v32 #whisper #nemotron_cascade_2

  25. Как я запускал Qwen 3.5 на Mac: бенчмарк 8 локальных LLM-серверов. Кто быстрее?

    Взял MacBook Pro M2 Max, 64GB, и задал простой вопрос: какой MLX-сервер реально готов держать Qwen 3.5 35B как локальный API для команды? Оказалось - серверов восемь, каждый в README обещает «blazing fast», а по факту между ними пропасть. Написал харнесс на Python, прогнал пять итераций на восьми промтах - от AIME до 52k токенов. Single-user тройка идёт ноздря в ноздрю. Но стоит пустить два запроса параллельно - и четыре фреймворка из шести откатываются в очередь, один деградирует до 0.85×, и только один выдаёт честные 2.17×. По дороге всплыли квадратичный attention в 2026 году, фантомные 14 000 tokens/sec из-за одной строчки в SSE-парсере и зомби-процесс на 20GB RAM, про который молчат все README. Внутри - графики, таблица «что выбрать под ваш сценарий» и репозиторий, чтобы повторить у себя.

    habr.com/ru/articles/1024880/

    #llm #Qwen3535BA3B #qwen35 #mlx #mac

  26. Как я запускал Qwen 3.5 на Mac: бенчмарк 8 локальных LLM-серверов. Кто быстрее?

    Взял MacBook Pro M2 Max, 64GB, и задал простой вопрос: какой MLX-сервер реально готов держать Qwen 3.5 35B как локальный API для команды? Оказалось - серверов восемь, каждый в README обещает «blazing fast», а по факту между ними пропасть. Написал харнесс на Python, прогнал пять итераций на восьми промтах - от AIME до 52k токенов. Single-user тройка идёт ноздря в ноздрю. Но стоит пустить два запроса параллельно - и четыре фреймворка из шести откатываются в очередь, один деградирует до 0.85×, и только один выдаёт честные 2.17×. По дороге всплыли квадратичный attention в 2026 году, фантомные 14 000 tokens/sec из-за одной строчки в SSE-парсере и зомби-процесс на 20GB RAM, про который молчат все README. Внутри - графики, таблица «что выбрать под ваш сценарий» и репозиторий, чтобы повторить у себя.

    habr.com/ru/articles/1024880/

    #llm #Qwen3535BA3B #qwen35 #mlx #mac

  27. Как мы перестали мерить качество ответов RAG-поиска «на глаз» и начали нормально сравнивать

    Если вы делаете RAG-поиск по документации или базе знаний, то рано или поздно упираетесь в проблему: хорошо найти — это еще не хорошо ответить. База знаний, RAG, найденные чанки, LLM строит ответ. Но пользователь не знает ни про DCG, ни про Recall@10, ни про чанки вообще. Он видит только то, что написано в итоговом ответе. А проблемы начинаются именно здесь: модель может что-то проигнорировать, ответить на другом языке, добавить что-то от себя или выдать уверенный текст с иероглифами посередине. В прошлой статье мы разбирали, как улучшали сам retrieval: чанкование, метаданные, гибридный поиск, реранкинг. Но после того как с поиском более-менее разобрались, встал другой вопрос — как вообще понять, хороший ли ответ получает пользователь? Привет, меня зовут Дима, я делаю ИИ-функции в

    habr.com/ru/companies/gram_ax/

    #RAG #AI #LLM #Qwen35 #Gemma_4 #gemma_3 #бенчмарк

  28. Как мы перестали мерить качество ответов RAG-поиска «на глаз» и начали нормально сравнивать

    Если вы делаете RAG-поиск по документации или базе знаний, то рано или поздно упираетесь в проблему: хорошо найти — это еще не хорошо ответить. База знаний, RAG, найденные чанки, LLM строит ответ. Но пользователь не знает ни про DCG, ни про Recall@10, ни про чанки вообще. Он видит только то, что написано в итоговом ответе. А проблемы начинаются именно здесь: модель может что-то проигнорировать, ответить на другом языке, добавить что-то от себя или выдать уверенный текст с иероглифами посередине. В прошлой статье мы разбирали, как улучшали сам retrieval: чанкование, метаданные, гибридный поиск, реранкинг. Но после того как с поиском более-менее разобрались, встал другой вопрос — как вообще понять, хороший ли ответ получает пользователь? Привет, меня зовут Дима, я делаю ИИ-функции в

    habr.com/ru/companies/gram_ax/

    #RAG #AI #LLM #Qwen35 #Gemma_4 #gemma_3 #бенчмарк

  29. I put 26.04 on a 2019 (7,1) with a 32GB HBM2 Vega II GPU. It's surprisingly AWESOME for both gaming (current-gen AAA GOG games through Heroic Launcher run very very well) AND quite awesome at local AI. getting nearly 80 tokens per second was really unexpected for this nearly obsolete box. All in all, silent and pretty good. Definitely runs better than MacOS.

  30. I put #Ubuntu 26.04 on a 2019 #MacPro (7,1) with a 32GB HBM2 Vega II GPU. It's surprisingly AWESOME for both gaming (current-gen AAA GOG games through Heroic Launcher run very very well) AND quite awesome at local AI. #Qwen35 getting nearly 80 tokens per second was really unexpected for this nearly obsolete box. All in all, silent and pretty good. Definitely runs better than MacOS.

  31. I put #Ubuntu 26.04 on a 2019 #MacPro (7,1) with a 32GB HBM2 Vega II GPU. It's surprisingly AWESOME for both gaming (current-gen AAA GOG games through Heroic Launcher run very very well) AND quite awesome at local AI. #Qwen35 getting nearly 80 tokens per second was really unexpected for this nearly obsolete box. All in all, silent and pretty good. Definitely runs better than MacOS.

  32. I put #Ubuntu 26.04 on a 2019 #MacPro (7,1) with a 32GB HBM2 Vega II GPU. It's surprisingly AWESOME for both gaming (current-gen AAA GOG games through Heroic Launcher run very very well) AND quite awesome at local AI. #Qwen35 getting nearly 80 tokens per second was really unexpected for this nearly obsolete box. All in all, silent and pretty good. Definitely runs better than MacOS.

  33. Иллюзия логики: как я доказал, что LLM-агенты игнорируют факты, и почему Chain-of-Thought делает только хуже

    Сейчас каждый второй стартап пилит ИИ-агентов. Мы оборачиваем LLM в цикл Промпт -> Вызов инструмента -> Ответ и ждем, что нейросеть сама расследует инцидент, найдет баг или напишет фичу. Но на практике автономные агенты часто ходят по кругу, игнорируют явные ошибки и «влюбляются» в свою первую догадку. Индустрия пытается лечить это костылями: наращивает контекст до миллионов токенов или заставляет модель «подумать шаг за шагом» (Chain-of-Thought). Я решил проверить эту архитектуру на прочность. Собрал локальный измерительный стенд LOCK-R, вооружился Теоремой Байеса и поймал современные LLM за руку. В этой статье я математически докажу, почему одиночные агенты структурно уязвимы, как токены размышлений заставляют их врать самим себе еще искуснее, и почему паттерн «Слепого Судьи» - это единственный способ вылечить AI от предвзятости. Тестируем на локальной Qwen-9B и фронтирной GPT-5.4.

    habr.com/ru/articles/1020016/

    #llm #ai_agents #rag #machine_learning #архитектура #chainofthought #теорема_байеса #gpt54 #qwen35 #бенчмарк

  34. Иллюзия логики: как я доказал, что LLM-агенты игнорируют факты, и почему Chain-of-Thought делает только хуже

    Сейчас каждый второй стартап пилит ИИ-агентов. Мы оборачиваем LLM в цикл Промпт -> Вызов инструмента -> Ответ и ждем, что нейросеть сама расследует инцидент, найдет баг или напишет фичу. Но на практике автономные агенты часто ходят по кругу, игнорируют явные ошибки и «влюбляются» в свою первую догадку. Индустрия пытается лечить это костылями: наращивает контекст до миллионов токенов или заставляет модель «подумать шаг за шагом» (Chain-of-Thought). Я решил проверить эту архитектуру на прочность. Собрал локальный измерительный стенд LOCK-R, вооружился Теоремой Байеса и поймал современные LLM за руку. В этой статье я математически докажу, почему одиночные агенты структурно уязвимы, как токены размышлений заставляют их врать самим себе еще искуснее, и почему паттерн «Слепого Судьи» - это единственный способ вылечить AI от предвзятости. Тестируем на локальной Qwen-9B и фронтирной GPT-5.4.

    habr.com/ru/articles/1020016/

    #llm #ai_agents #rag #machine_learning #архитектура #chainofthought #теорема_байеса #gpt54 #qwen35 #бенчмарк

  35. So I've found that #Qwen35's training data knows everything about #Artemis up to this launch, #Artemis2. Knew the astronaut names, etc.

    I only had vague notions about future missions, so I asked about it. And it mentioned the "Lunar Gateway". I looked that up on Wikipedia and it was quite accurate. However... it had no way to know the Gateway (an orbiting lunar support station) was axed by the Trump administration in favor of going directly towards building a lunar base.

    Sounds to me like some orange baby said _"I want my admin to put a base on the moon, not just another lame ISS! DO IT OR YOU DON'T GET FUNDING!!"_ 🤷

    But I'm open to different interpretations, of course. I'm just skeptical and ignorant of the actual science needs.

    en.wikipedia.org/wiki/Lunar_Ga

    > On July 4, 2025, President Donald Trump signed the One Big Beautiful Bill Act into law, allocating $2.6 billion for the program and requiring at least $750 million annually from FY 2026 through FY 2028.
    >
    > In early 2026, reports indicated that references to the station had been removed from congressional funding legislation. On February 26, 2026, reporting suggested that NASA Administrator Jared Isaacman was considering restructuring the program toward a lunar surface base effort in Houston.
    >
    > In March 2026, NASA announced it would no longer build the station and would instead focus on a lunar surface base between 2029 and 2036, repurposing Gateway hardware and partner contributions where possible. Carlos Garcia-Galan, NASA's program manager for the Lunar Gateway, was reassigned to lead the surface base effort but stated that a lunar orbiting outpost "has value in our overall exploration goals" and that NASA may consider it later, but that the agency is now focused on the surface.

    #nasa #llm #localLLM

  36. So I've found that #Qwen35's training data knows everything about #Artemis up to this launch, #Artemis2. Knew the astronaut names, etc.

    I only had vague notions about future missions, so I asked about it. And it mentioned the "Lunar Gateway". I looked that up on Wikipedia and it was quite accurate. However... it had no way to know the Gateway (an orbiting lunar support station) was axed by the Trump administration in favor of going directly towards building a lunar base.

    Sounds to me like some orange baby said _"I want my admin to put a base on the moon, not just another lame ISS! DO IT OR YOU DON'T GET FUNDING!!"_ 🤷

    But I'm open to different interpretations, of course. I'm just skeptical and ignorant of the actual science needs.

    en.wikipedia.org/wiki/Lunar_Ga

    > On July 4, 2025, President Donald Trump signed the One Big Beautiful Bill Act into law, allocating $2.6 billion for the program and requiring at least $750 million annually from FY 2026 through FY 2028.
    >
    > In early 2026, reports indicated that references to the station had been removed from congressional funding legislation. On February 26, 2026, reporting suggested that NASA Administrator Jared Isaacman was considering restructuring the program toward a lunar surface base effort in Houston.
    >
    > In March 2026, NASA announced it would no longer build the station and would instead focus on a lunar surface base between 2029 and 2036, repurposing Gateway hardware and partner contributions where possible. Carlos Garcia-Galan, NASA's program manager for the Lunar Gateway, was reassigned to lead the surface base effort but stated that a lunar orbiting outpost "has value in our overall exploration goals" and that NASA may consider it later, but that the agency is now focused on the surface.

    #nasa #llm #localLLM

  37. So I've found that #Qwen35's training data knows everything about #Artemis up to this launch, #Artemis2. Knew the astronaut names, etc.

    I only had vague notions about future missions, so I asked about it. And it mentioned the "Lunar Gateway". I looked that up on Wikipedia and it was quite accurate. However... it had no way to know the Gateway (an orbiting lunar support station) was axed by the Trump administration in favor of going directly towards building a lunar base.

    Sounds to me like some orange baby said _"I want my admin to put a base on the moon, not just another lame ISS! DO IT OR YOU DON'T GET FUNDING!!"_ 🤷

    But I'm open to different interpretations, of course. I'm just skeptical and ignorant of the actual science needs.

    en.wikipedia.org/wiki/Lunar_Ga

    > On July 4, 2025, President Donald Trump signed the One Big Beautiful Bill Act into law, allocating $2.6 billion for the program and requiring at least $750 million annually from FY 2026 through FY 2028.
    >
    > In early 2026, reports indicated that references to the station had been removed from congressional funding legislation. On February 26, 2026, reporting suggested that NASA Administrator Jared Isaacman was considering restructuring the program toward a lunar surface base effort in Houston.
    >
    > In March 2026, NASA announced it would no longer build the station and would instead focus on a lunar surface base between 2029 and 2036, repurposing Gateway hardware and partner contributions where possible. Carlos Garcia-Galan, NASA's program manager for the Lunar Gateway, was reassigned to lead the surface base effort but stated that a lunar orbiting outpost "has value in our overall exploration goals" and that NASA may consider it later, but that the agency is now focused on the surface.

    #nasa #llm #localLLM

  38. So I've found that #Qwen35's training data knows everything about #Artemis up to this launch, #Artemis2. Knew the astronaut names, etc.

    I only had vague notions about future missions, so I asked about it. And it mentioned the "Lunar Gateway". I looked that up on Wikipedia and it was quite accurate. However... it had no way to know the Gateway (an orbiting lunar support station) was axed by the Trump administration in favor of going directly towards building a lunar base.

    Sounds to me like some orange baby said _"I want my admin to put a base on the moon, not just another lame ISS! DO IT OR YOU DON'T GET FUNDING!!"_ 🤷

    But I'm open to different interpretations, of course. I'm just skeptical and ignorant of the actual science needs.

    en.wikipedia.org/wiki/Lunar_Ga

    > On July 4, 2025, President Donald Trump signed the One Big Beautiful Bill Act into law, allocating $2.6 billion for the program and requiring at least $750 million annually from FY 2026 through FY 2028.
    >
    > In early 2026, reports indicated that references to the station had been removed from congressional funding legislation. On February 26, 2026, reporting suggested that NASA Administrator Jared Isaacman was considering restructuring the program toward a lunar surface base effort in Houston.
    >
    > In March 2026, NASA announced it would no longer build the station and would instead focus on a lunar surface base between 2029 and 2036, repurposing Gateway hardware and partner contributions where possible. Carlos Garcia-Galan, NASA's program manager for the Lunar Gateway, was reassigned to lead the surface base effort but stated that a lunar orbiting outpost "has value in our overall exploration goals" and that NASA may consider it later, but that the agency is now focused on the surface.

    #nasa #llm #localLLM

  39. LLMArena and specifically their coding leaderboard, where 27b Qwen model is 20 positions higher than 675b Mistral model shows really good, how slow Mistral is and how much they are lagging behine Chinese open source competition, not even mentioning American SOTA models.

    #mistral #mistralai #qwen #qwen35 #ai #artificialintelligence

  40. LLMArena and specifically their coding leaderboard, where 27b Qwen model is 20 positions higher than 675b Mistral model shows really good, how slow Mistral is and how much they are lagging behine Chinese open source competition, not even mentioning American SOTA models.

    #mistral #mistralai #qwen #qwen35 #ai #artificialintelligence

  41. LLMArena and specifically their coding leaderboard, where 27b Qwen model is 20 positions higher than 675b Mistral model shows really good, how slow Mistral is and how much they are lagging behine Chinese open source competition, not even mentioning American SOTA models.

    #mistral #mistralai #qwen #qwen35 #ai #artificialintelligence

  42. LLMArena and specifically their coding leaderboard, where 27b Qwen model is 20 positions higher than 675b Mistral model shows really good, how slow Mistral is and how much they are lagging behine Chinese open source competition, not even mentioning American SOTA models.

    #mistral #mistralai #qwen #qwen35 #ai #artificialintelligence

  43. La precedente esperienza con Qwen3.5 non aveva dato i risultati sperati. Nonostante ore di lavoro e feedback continui, il modello non è mai riuscito a produrre un’applicazione funzionante: regressioni cicliche ed errori difficilmente superabili con le capacità dello strumento hanno bloccato ogni progresso.

    Ho voluto quindi riprovare con Nemotron-Cascade-2, ma le sue richieste hardware si […]

    #agenticAi #ai #claudeCode #nemotron #openrouter #qwen35 https://www.b0sh.net/2026/03/nemotron-3-super-vs-qwen3-5-costruire-unapp-con-lai-senza-scrivere-codice/
  44. La precedente esperienza con Qwen3.5 non aveva dato i risultati sperati. Nonostante ore di lavoro e feedback continui, il modello non è mai riuscito a produrre un’applicazione funzionante: regressioni cicliche ed errori difficilmente superabili con le capacità dello strumento hanno bloccato ogni progresso.

    Ho voluto quindi riprovare con Nemotron-Cascade-2, ma le sue richieste hardware si […]

    #agenticAi #ai #claudeCode #nemotron #openrouter #qwen35 https://www.b0sh.net/2026/03/nemotron-3-super-vs-qwen3-5-costruire-unapp-con-lai-senza-scrivere-codice/
  45. La precedente esperienza con Qwen3.5 non aveva dato i risultati sperati. Nonostante ore di lavoro e feedback continui, il modello non è mai riuscito a produrre un’applicazione funzionante: regressioni cicliche ed errori difficilmente superabili con le capacità dello strumento hanno bloccato ogni progresso.

    Ho voluto quindi riprovare con Nemotron-Cascade-2, ma le sue richieste hardware si […]

    #agenticAi #ai #claudeCode #nemotron #openrouter #qwen35 https://www.b0sh.net/2026/03/nemotron-3-super-vs-qwen3-5-costruire-unapp-con-lai-senza-scrivere-codice/
  46. La precedente esperienza con Qwen3.5 non aveva dato i risultati sperati. Nonostante ore di lavoro e feedback continui, il modello non è mai riuscito a produrre un’applicazione funzionante: regressioni cicliche ed errori difficilmente superabili con le capacità dello strumento hanno bloccato ogni progresso.

    Ho voluto quindi riprovare con Nemotron-Cascade-2, ma le sue richieste hardware si […]

    #agenticAi #ai #claudeCode #nemotron #openrouter #qwen35 https://www.b0sh.net/2026/03/nemotron-3-super-vs-qwen3-5-costruire-unapp-con-lai-senza-scrivere-codice/
  47. It's amusing that #Qwen35 has been particularly sensitive to dates set in "the future" of it's training set. It even called a bunch of recent MCU movies referenced in one particular article as "imaginary" and "just for fun". 😏

    #llm #ai

  48. It's amusing that #Qwen35 has been particularly sensitive to dates set in "the future" of it's training set. It even called a bunch of recent MCU movies referenced in one particular article as "imaginary" and "just for fun". 😏

    #llm #ai

  49. It's amusing that #Qwen35 has been particularly sensitive to dates set in "the future" of it's training set. It even called a bunch of recent MCU movies referenced in one particular article as "imaginary" and "just for fun". 😏

    #llm #ai

  50. It's amusing that #Qwen35 has been particularly sensitive to dates set in "the future" of it's training set. It even called a bunch of recent MCU movies referenced in one particular article as "imaginary" and "just for fun". 😏

    #llm #ai

  51. 好明顯今次release的Qwen3.5是有針對特定的硬件,例如8G VRAM 的顯卡,Amd Strix Halo CPU, Apple M chips

    #qwen35

  52. On X , people saw a five word post from Junyang Lin, the man who built Qwen from the ground up: “bye my beloved qwen.”

    That was it. No explanation, just a goodbye.

    Within hours the replies were flooding in. Developers, researchers, open source contributors all asking the same thing, what just happened?

    #qwen35 #alibaba #ai #qwen

    firethering.com/qwen-core-team

  53. firethering.com/qwen3-5-4b-loc
    Alibaba just dropped #Qwen35 and the 4B version is the one worth paying attention to. It thinks before it answers, reads images and video, handles 201 languages, and sits on a context window of 262,144 tokens, longer than most models ten times its size. #opensource

  54. Из коробки не работает: запускаем свежие большие LLM

    В последнее время открытых моделей сверхбольшого размера развелось неимоверное количество, даже не просто моделей, а производителей. Вариации GLM, Kimi, DeepSeek занимают по нескольку строк в топ 5-10-20. Понадобилось перебрать основные LLM для тестов и выбора "рабочей лошадки", для чего пришлось немного пошуршать в интернетах. Оставлю в качестве памятки, вдруг кому-то окажется полезным. Всё делалось на базе образов vllm-openai, платформ B200/H200 и дров 590.48.01. На момент начала экспериментов - примерно пару недель тому назад - версии vllm 0.16 ещё не было, но, как выяснилось в итоге, это не сильно повлияло на ситуацию. Основные костыли остались теми же самыми. Разве что кастомизация образа не для каждой модели нужна теперь. В целом там, понятное дело, никакого RocketScience нету (особенно после того, как почитаешь китайские форумы в поисках нюансов). Но если бы кто-то посидел заранее и собрал советы в одном месте - жизнь была бы немного проще )) поэтому делюсь. Итак, поехали.

    habr.com/ru/articles/1006202/

    #KimiK25 #DeepSeekv32 #GLM5 #Qwen35 #vllm #B200 #H200

  55. Из коробки не работает: запускаем свежие большие LLM

    В последнее время открытых моделей сверхбольшого размера развелось неимоверное количество, даже не просто моделей, а производителей. Вариации GLM, Kimi, DeepSeek занимают по нескольку строк в топ 5-10-20. Понадобилось перебрать основные LLM для тестов и выбора "рабочей лошадки", для чего пришлось немного пошуршать в интернетах. Оставлю в качестве памятки, вдруг кому-то окажется полезным. Всё делалось на базе образов vllm-openai, платформ B200/H200 и дров 590.48.01. На момент начала экспериментов - примерно пару недель тому назад - версии vllm 0.16 ещё не было, но, как выяснилось в итоге, это не сильно повлияло на ситуацию. Основные костыли остались теми же самыми. Разве что кастомизация образа не для каждой модели нужна теперь. В целом там, понятное дело, никакого RocketScience нету (особенно после того, как почитаешь китайские форумы в поисках нюансов). Но если бы кто-то посидел заранее и собрал советы в одном месте - жизнь была бы немного проще )) поэтому делюсь. Итак, поехали.

    habr.com/ru/articles/1006202/

    #KimiK25 #DeepSeekv32 #GLM5 #Qwen35 #vllm #B200 #H200