home.social

#qwen — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #qwen, aggregated by home.social.

fetched live
  1. 🧪 LLM Benchmark Showdown: 5 lokale Ollama-Modelle im Vergleich

    Getestet auf derselben Hardware (#gmktecevo2 ):
    (100 Samples) — Math
    (100/Kategorie) — Function Calling
    + (50) — Python Coding
    + (20) — Python Coding

    📊 Ergebnisse (Accuracy / Output TK/s / VRAM):

    **qwen3.8:27b**
    GSM8K 82% | BFCL 91.5% | MBPP+ 100% | HE+ 100%
    ⚡ 25.5 TK/s | 💾 18 GB VRAM

    **qwen3.6:27b**
    GSM8K 83% | BFCL 93% | MBPP+ 98% | HE+ 75%
    ⚡ 12.7 TK/s | 💾 33 GB VRAM

    **qwen3.6:35b**
    GSM8K 84% | BFCL 90% | MBPP+ 98% | HE+ 55%
    ⚡ 61.8 TK/s | 💾 27 GB VRAM

    **ornith-1.5:35b**
    GSM8K 75% | BFCL 92.5% | MBPP+ 78% | HE+ 0%
    ⚡ 63.6 TK/s | 💾 26 GB VRAM

    **nemotron-3.5-lightning:30b**
    GSM8K 59% | BFCL 74% | MBPP+ 94% | HE+ 0%
    ⚡ 91.9 TK/s | 💾 26 GB VRAM

    🏆 Fazit:

    qwen3.8:27b ist der klare Sieger — als einziges Modell 100% bei beiden Coding-Benchmarks, bei GSM8K/BFCL gleichauf mit den anderen Qwen-Modellen. Bei 25.5 TK/s und nur 18 GB VRAM das beste Qualität/Speed/Effizienz-Verhältnis.

    qwen3.6:27b ist qualitativ nah dran (BFCL sogar 93%), aber mit 12.7 TK/s unerträglich langsam und frisst 33 GB VRAM — fast 2× so viel wie qwen3.8 bei halber Speed.

    qwen3.6:35b ist mit 61.8 TK/s 2.4× schneller als qwen3.8, aber HE+ nur 55% (vs 100%). Trading Code-Qualität für Speed.

    ornith-1.5:35b und nemotron-3.5-lightning:30b fallen bei Coding komplett durch (HE+ 0%), sind aber die schnellsten Modelle im Feld (64 / 92 TK/s).

    💡 TK/s = generierte Tokens/Sekunde (Warm-Run, ollama --verbose).
    💾 VRAM = GPU-Speicher bei max context (262K bzw. 1M bei nemotron).

  2. Suite de mes pérégrinations uncensored : par curiosité et affinité, j’ai soumis Qwen3.8-27B-Uncensored (Orca) au CTIBench 2024, histoire de voir ce qu’elle avait vraiment dans le ventre côté CTI.
    👇
    github.com/maveryn/cti-bench

    Version F16 full, 27B, sur H100.
    2 500 questions CTI-MCQ.

    Résultat : 70,64 %.

    À titre de repère, dans le papier CTIBench original :

    • GPT-4 : 71,0 %
    • Llama 3 70B : 65,72 %
    • Gemini 1.5 : 65,44 %
    • Llama 3 8B : 61,32 %

    Donc ce petit 27B uncensored finit à 0,36 point du GPT-4 testé en 2024.

    Évidemment, comparaison historique à prendre avec les pincettes habituelles : benchmark public depuis 2024, modèles 2026, contamination impossible à exclure, etc...

    Mais quand même… pas mal du tout.
    #Qwen

  3. Suite de mes pérégrinations uncensored : par curiosité et affinité, j’ai soumis Qwen3.8-27B-Uncensored (Orca) au CTIBench 2024, histoire de voir ce qu’elle avait vraiment dans le ventre côté CTI.
    👇
    github.com/maveryn/cti-bench

    Version F16 full, 27B, sur H100.
    2 500 questions CTI-MCQ.

    Résultat : 70,64 %.

    À titre de repère, dans le papier CTIBench original :

    • GPT-4 : 71,0 %
    • Llama 3 70B : 65,72 %
    • Gemini 1.5 : 65,44 %
    • Llama 3 8B : 61,32 %

    Donc ce petit 27B uncensored finit à 0,36 point du GPT-4 testé en 2024.

    Évidemment, comparaison historique à prendre avec les pincettes habituelles : benchmark public depuis 2024, modèles 2026, contamination impossible à exclure, etc...

    Mais quand même… pas mal du tout.
    #Qwen

  4. Suite de mes pérégrinations uncensored : par curiosité et affinité, j’ai soumis Qwen3.8-27B-Uncensored (Orca) au CTIBench 2024, histoire de voir ce qu’elle avait vraiment dans le ventre côté CTI.
    👇
    github.com/maveryn/cti-bench

    Version F16 full, 27B, sur H100.
    2 500 questions CTI-MCQ.

    Résultat : 70,64 %.

    À titre de repère, dans le papier CTIBench original :

    • GPT-4 : 71,0 %
    • Llama 3 70B : 65,72 %
    • Gemini 1.5 : 65,44 %
    • Llama 3 8B : 61,32 %

    Donc ce petit 27B uncensored finit à 0,36 point du GPT-4 testé en 2024.

    Évidemment, comparaison historique à prendre avec les pincettes habituelles : benchmark public depuis 2024, modèles 2026, contamination impossible à exclure, etc...

    Mais quand même… pas mal du tout.
    #Qwen

  5. Suite de mes pérégrinations uncensored : par curiosité et affinité, j’ai soumis Qwen3.8-27B-Uncensored (Orca) au CTIBench 2024, histoire de voir ce qu’elle avait vraiment dans le ventre côté CTI.
    👇
    github.com/maveryn/cti-bench

    Version F16 full, 27B, sur H100.
    2 500 questions CTI-MCQ.

    Résultat : 70,64 %.

    À titre de repère, dans le papier CTIBench original :

    • GPT-4 : 71,0 %
    • Llama 3 70B : 65,72 %
    • Gemini 1.5 : 65,44 %
    • Llama 3 8B : 61,32 %

    Donc ce petit 27B uncensored finit à 0,36 point du GPT-4 testé en 2024.

    Évidemment, comparaison historique à prendre avec les pincettes habituelles : benchmark public depuis 2024, modèles 2026, contamination impossible à exclure, etc...

    Mais quand même… pas mal du tout.
    #Qwen

  6. CW: AI

    Решил затестить Qwen3.8-27B-GGUF на своей 9060XT. Взял 3 бита XL вроде.

    Аухел, что оно смогло наклепать простенький лендинг и всё это чисто на моём железе, жест

    #qwen #qwen3

  7. CW: AI

    Решил затестить Qwen3.8-27B-GGUF на своей 9060XT. Взял 3 бита XL вроде.

    Аухел, что оно смогло наклепать простенький лендинг и всё это чисто на моём железе, жест

    #qwen #qwen3

  8. CW: AI

    Решил затестить Qwen3.8-27B-GGUF на своей 9060XT. Взял 3 бита XL вроде.

    Аухел, что оно смогло наклепать простенький лендинг и всё это чисто на моём железе, жест

    #qwen #qwen3

  9. RT @cgtwts: Wir sind möglicherweise viel näher daran, Modelle lokal auszuführen, als jeder erwartet hat. FreeToken von UC Berkeley und MIT ermöglicht: -DeepSeek-V4-Flash 284B mit 25 Token pro Sekunde auf einer einzelnen RTX 5090, - während ein 8GB RTX 4060 Laptop bei Qwen3.6-35B 39 Token pro Sekunde erreicht. Und es ist 2–4x schneller als Ollama auf Consumer-GPUs.

    mehr auf Arint.info

    #DeepSeek #KI #LokaleModelle #OpenSource #Qwen #RTX5090 #arint_info

    https://x.com/cgtwts/status/2091149529274085436

  10. RT @cgtwts: Wir sind möglicherweise viel näher daran, Modelle lokal auszuführen, als jeder erwartet hat. FreeToken von UC Berkeley und MIT ermöglicht: -DeepSeek-V4-Flash 284B mit 25 Token pro Sekunde auf einer einzelnen RTX 5090, - während ein 8GB RTX 4060 Laptop bei Qwen3.6-35B 39 Token pro Sekunde erreicht. Und es ist 2–4x schneller als Ollama auf Consumer-GPUs.

    mehr auf Arint.info

    #DeepSeek #KI #LokaleModelle #OpenSource #Qwen #RTX5090 #arint_info

    https://x.com/cgtwts/status/2091149529274085436

  11. RT @cgtwts: Wir sind möglicherweise viel näher daran, Modelle lokal auszuführen, als jeder erwartet hat. FreeToken von UC Berkeley und MIT ermöglicht: -DeepSeek-V4-Flash 284B mit 25 Token pro Sekunde auf einer einzelnen RTX 5090, - während ein 8GB RTX 4060 Laptop bei Qwen3.6-35B 39 Token pro Sekunde erreicht. Und es ist 2–4x schneller als Ollama auf Consumer-GPUs.

    mehr auf Arint.info

    #DeepSeek #KI #LokaleModelle #OpenSource #Qwen #RTX5090 #arint_info

    https://x.com/cgtwts/status/2091149529274085436

  12. Building a local AI storyteller - Part II - When the model is not enough
    A blog by Justus

    In the first part of this series, I described the most important lesson I learned while building a local AI storyteller: A good LLM application is not a clever prompt. It is software architecture around a probabilistic component. That sounds reassuringly architectural. It also leaves one...

    #dev #softwaredevelopment #java #ai #llm #localmodels #gemma #qwen #softwarearchitecture

    jdriven.com/blog/2026/08/llm-s

  13. Building a local AI storyteller - Part II - When the model is not enough
    A blog by Justus

    In the first part of this series, I described the most important lesson I learned while building a local AI storyteller: A good LLM application is not a clever prompt. It is software architecture around a probabilistic component. That sounds reassuringly architectural. It also leaves one...

    #dev #softwaredevelopment #java #ai #llm #localmodels #gemma #qwen #softwarearchitecture

    jdriven.com/blog/2026/08/llm-s

  14. Tried the same as above with muse-glimmer:30b. The output is **much** better. Only 67 lines of python, no errors, a proper board and straight-forwards input. I'm guessing Qwen 3.8 is benchmark optimized.

    #AI #LLM #Qwen #Muse #LocalLLM

  15. RT @LiuVaayne: Auf meinem Computer verwende ich vorübergehend Ornith-1.5 35B A3B anstelle von Qwen 3.8 27B. Qwen wurde lange optimiert, erreicht aber nur 20 tok/s, während Ornith ohne Anpassungen 100 tok/s liefert und somit für viele Aufgaben genutzt werden kann. Die aktuell auf dem Computer eingesetzten Modelle sind: - Ornith-1.5 35B A3B für den täglichen Gebrauch - Hy-MT2-1.8B für Übersetzungen - Qwen3-Embedding-4B für Embeddings - Unlimited-OCR für OCR-Aufgaben

    mehr auf Arint.info

    #Embedding #KIModelle #MaschinelleÜbersetzung #OCR #Ornith #Qwen #arint_info

    https://x.com/LiuVaayne/status/2090619913279021231

  16. RT @LiuVaayne: Auf meinem Computer verwende ich vorübergehend Ornith-1.5 35B A3B anstelle von Qwen 3.8 27B. Qwen wurde lange optimiert, erreicht aber nur 20 tok/s, während Ornith ohne Anpassungen 100 tok/s liefert und somit für viele Aufgaben genutzt werden kann. Die aktuell auf dem Computer eingesetzten Modelle sind: - Ornith-1.5 35B A3B für den täglichen Gebrauch - Hy-MT2-1.8B für Übersetzungen - Qwen3-Embedding-4B für Embeddings - Unlimited-OCR für OCR-Aufgaben

    mehr auf Arint.info

    #Embedding #KIModelle #MaschinelleÜbersetzung #OCR #Ornith #Qwen #arint_info

    https://x.com/LiuVaayne/status/2090619913279021231

  17. RT @LiuVaayne: Auf meinem Computer verwende ich vorübergehend Ornith-1.5 35B A3B anstelle von Qwen 3.8 27B. Qwen wurde lange optimiert, erreicht aber nur 20 tok/s, während Ornith ohne Anpassungen 100 tok/s liefert und somit für viele Aufgaben genutzt werden kann. Die aktuell auf dem Computer eingesetzten Modelle sind: - Ornith-1.5 35B A3B für den täglichen Gebrauch - Hy-MT2-1.8B für Übersetzungen - Qwen3-Embedding-4B für Embeddings - Unlimited-OCR für OCR-Aufgaben

    mehr auf Arint.info

    #Embedding #KIModelle #MaschinelleÜbersetzung #OCR #Ornith #Qwen #arint_info

    https://x.com/LiuVaayne/status/2090619913279021231

  18. More #LocalLLM test results:

    #Qwen 3.6 #35B A3B #mmproj on description of foto collection. 2x12GB VRAM (at ~50-60t/s):

    * ca 8800 fotos, 88 GB
    * almost flawlessly accurate scene descriptions
    * perfectly usable for search/retrieval by keywords

    Attached is 1 example, describing the "Papiergießkanne"

    wow.
    Ran at ~250 W for ~10h.

    I get ~25kWh on a sunny day from my roof. 🌄

  19. More #LocalLLM test results:

    #Qwen 3.6 #35B A3B #mmproj on description of foto collection. 2x12GB VRAM (at ~50-60t/s):

    * ca 8800 fotos, 88 GB
    * almost flawlessly accurate scene descriptions
    * perfectly usable for search/retrieval by keywords

    Attached is 1 example, describing the "Papiergießkanne"

    wow.
    Ran at ~250 W for ~10h.

    I get ~25kWh on a sunny day from my roof. 🌄

  20. More #LocalLLM test results:

    #Qwen 3.6 #35B A3B #mmproj on description of foto collection. 2x12GB VRAM (at ~50-60t/s):

    * ca 8800 fotos, 88 GB
    * almost flawlessly accurate scene descriptions
    * perfectly usable for search/retrieval by keywords

    Attached is 1 example, describing the "Papiergießkanne"

    wow.
    Ran at ~250 W for ~10h.

    I get ~25kWh on a sunny day from my roof. 🌄

  21. OK, I ran my usual quick coding test of having #qwen 3.8 whip up a console tic-tac-toe game in python. I have to say, that I am *not* impressed. The file is around 178 lines long, but the first pass had a syntax error, which I had the #LLM fix. When the game did run, it did not display the board in any sensible way at all. Many smaller models have done much better than this in the past and the best local model for coding is still Qwen3-coder:30b which was *much* faster with better output.

    #AI

  22. So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.

    Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.

    Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark

  23. So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.

    Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.

    Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark

  24. So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.

    Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.

    Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark

  25. So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.

    Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.

    Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark

  26. So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.

    Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.

    Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark

  27. Свой инференс для 25 разработчиков: 452:1, KV‑пул и почему это не экономит денег

    Для тех, кто держит или собирается держать LLM внутри контура: тимлидов, DevOps, архитекторов. Здесь конфиги, цифры и грабли, а не введение в трансформеры. Что вы унесёте: историю пяти последовательных конфигураций с тем, что каждая дала и чего стоила; рабочий набор флагов vLLM под одну карту Blackwell; три неочевидных бага и обходы; разбор реального счёта с отношением вход/выход 452:1. Главный вывод, если дальше читать некогда: на агентской нагрузке кэш префикса решает больше, чем выбор модели, размер карты и всё остальное вместе взятое. Мы шли к этому через четыре промежуточных стенда и потратили лишние месяцы, потому что не понимали, что именно меряем.

    habr.com/ru/articles/1073300/

    #vLLM #LiteLLM #prefix_caching #KVcache #локальные_LLM #инференс_LLM #Qwen #агентская_разработка #RTX_PRO_6000 #claude

  28. Свой инференс для 25 разработчиков: 452:1, KV‑пул и почему это не экономит денег

    Для тех, кто держит или собирается держать LLM внутри контура: тимлидов, DevOps, архитекторов. Здесь конфиги, цифры и грабли, а не введение в трансформеры. Что вы унесёте: историю пяти последовательных конфигураций с тем, что каждая дала и чего стоила; рабочий набор флагов vLLM под одну карту Blackwell; три неочевидных бага и обходы; разбор реального счёта с отношением вход/выход 452:1. Главный вывод, если дальше читать некогда: на агентской нагрузке кэш префикса решает больше, чем выбор модели, размер карты и всё остальное вместе взятое. Мы шли к этому через четыре промежуточных стенда и потратили лишние месяцы, потому что не понимали, что именно меряем.

    habr.com/ru/articles/1073300/

    #vLLM #LiteLLM #prefix_caching #KVcache #локальные_LLM #инференс_LLM #Qwen #агентская_разработка #RTX_PRO_6000 #claude

  29. Свой инференс для 25 разработчиков: 452:1, KV‑пул и почему это не экономит денег

    Для тех, кто держит или собирается держать LLM внутри контура: тимлидов, DevOps, архитекторов. Здесь конфиги, цифры и грабли, а не введение в трансформеры. Что вы унесёте: историю пяти последовательных конфигураций с тем, что каждая дала и чего стоила; рабочий набор флагов vLLM под одну карту Blackwell; три неочевидных бага и обходы; разбор реального счёта с отношением вход/выход 452:1. Главный вывод, если дальше читать некогда: на агентской нагрузке кэш префикса решает больше, чем выбор модели, размер карты и всё остальное вместе взятое. Мы шли к этому через четыре промежуточных стенда и потратили лишние месяцы, потому что не понимали, что именно меряем.

    habr.com/ru/articles/1073300/

    #vLLM #LiteLLM #prefix_caching #KVcache #локальные_LLM #инференс_LLM #Qwen #агентская_разработка #RTX_PRO_6000 #claude

  30. RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.

    mehr auf Arint.info

    #AI #KI #LLM #MachineLearning #Qwen #Tech #arint_info

    https://x.com/ivanfioravanti/status/2090100203537740099

  31. RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.

    mehr auf Arint.info

    #AI #KI #LLM #MachineLearning #Qwen #Tech #arint_info

    https://x.com/ivanfioravanti/status/2090100203537740099

  32. RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.

    mehr auf Arint.info

    #AI #KI #LLM #MachineLearning #Qwen #Tech #arint_info

    https://x.com/ivanfioravanti/status/2090100203537740099

  33. RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.

    mehr auf Arint.info

    #AI #KI #LLM #MachineLearning #Qwen #Tech #arint_info

    https://x.com/ivanfioravanti/status/2090100203537740099

  34. RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.

    mehr auf Arint.info

    #AI #KI #LLM #MachineLearning #Qwen #Tech #arint_info

    https://x.com/ivanfioravanti/status/2090100203537740099

  35. I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”

    Output speed:
    * muse-glimmer:30b
    * eval rate: 1.85 tokens/s

    * qwen3.8:27b
    * eval rate: 1.42 tokens/s

    Two observations:
    1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
    2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP).

    #AI #LLM #Qwen #LocalLLM #VibeCoding #Programming #Coding

  36. I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”

    Output speed:
    * muse-glimmer:30b
    * eval rate: 1.85 tokens/s

    * qwen3.8:27b
    * eval rate: 1.42 tokens/s

    Two observations:
    1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
    2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP).

    #AI #LLM #Qwen #LocalLLM #VibeCoding #Programming #Coding

  37. I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”

    Output speed:
    * muse-glimmer:30b
    * eval rate: 1.85 tokens/s

    * qwen3.8:27b
    * eval rate: 1.42 tokens/s

    Two observations:
    1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
    2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP).

    #AI #LLM #Qwen #LocalLLM #VibeCoding #Programming #Coding

  38. I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”

    Output speed:
    * muse-glimmer:30b
    * eval rate: 1.85 tokens/s

    * qwen3.8:27b
    * eval rate: 1.42 tokens/s

    Two observations:
    1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
    2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP).

    #AI #LLM #Qwen #LocalLLM #VibeCoding #Programming #Coding

  39. I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”

    Output speed:
    * muse-glimmer:30b
    * eval rate: 1.85 tokens/s

    * qwen3.8:27b
    * eval rate: 1.42 tokens/s

    Two observations:
    1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
    2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP).

    #AI #LLM #Qwen #LocalLLM #VibeCoding #Programming #Coding

  40. qwen3.8-27b:q4_K_M running on my RTX-4090 via very nicely. I couldn't make it run on the latest

    I'm tired of using as my primary agent. I'm trialing switching to with and using CC as a subagent.

  41. FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.

    Inference is just WAY too slow.

    I got the same work done with 3.6 MTP in less than 10% of the time.

    Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.

    I'll wait for a month until they come up with an update (like 3.6 MTP).

    Details matter, feel free to query me.

    #AI

  42. FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.

    Inference is just WAY too slow.

    I got the same work done with 3.6 MTP in less than 10% of the time.

    Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.

    I'll wait for a month until they come up with an update (like 3.6 MTP).

    Details matter, feel free to query me.

    #AI

  43. FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.

    Inference is just WAY too slow.

    I got the same work done with 3.6 MTP in less than 10% of the time.

    Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.

    I'll wait for a month until they come up with an update (like 3.6 MTP).

    Details matter, feel free to query me.

    #AI

  44. FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.

    Inference is just WAY too slow.

    I got the same work done with 3.6 MTP in less than 10% of the time.

    Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.

    I'll wait for a month until they come up with an update (like 3.6 MTP).

    Details matter, feel free to query me.

    #AI

  45. FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.

    Inference is just WAY too slow.

    I got the same work done with 3.6 MTP in less than 10% of the time.

    Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.

    I'll wait for a month until they come up with an update (like 3.6 MTP).

    Details matter, feel free to query me.

    #AI

  46. Мой опыт с Hermes Agent — ненависть, любовь, ненависть, любовь

    Началось все с установки. Я пошёл почти по самому простому пути - установил его на Mac, в Docker. Потому что это агент, который работает автономно и так же автономно может сделать атата: выполнить rm -rf или выбраться из клетки и начать всё взламывать :) В рамках настройки я сразу выдал доступ к части файлов только на чтение, и только к Obsidian - на чтение и запись (потому что писать он в данном случае должен), но файлы были под Git.

    habr.com/ru/articles/1072770/

    #hermes_agent #aiагенты #llm #локальный_llm #локальные_модели #qwen #agentic_workflows #ai_automation #workflow #go

  47. Мой опыт с Hermes Agent — ненависть, любовь, ненависть, любовь

    Началось все с установки. Я пошёл почти по самому простому пути - установил его на Mac, в Docker. Потому что это агент, который работает автономно и так же автономно может сделать атата: выполнить rm -rf или выбраться из клетки и начать всё взламывать :) В рамках настройки я сразу выдал доступ к части файлов только на чтение, и только к Obsidian - на чтение и запись (потому что писать он в данном случае должен), но файлы были под Git.

    habr.com/ru/articles/1072770/

    #hermes_agent #aiагенты #llm #локальный_llm #локальные_модели #qwen #agentic_workflows #ai_automation #workflow #go

  48. Testing #localLLM for #DLTP on my #longterm #preservation #object identifier schema of "CFIDs":

    Collision Friendly IDentifiers.

    A mere 30 MB list of auto-generated CFIDs was enough for #qwen to tell you this much about the test collection!!! 🤯 🤩 - this is powerful. Be careful.

    AND: My IDs work! so beautiful! #ahalodeck

    github.com/ArkThis/AHAlodeck/b