#qwen — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #qwen, aggregated by home.social.
-
🧪 LLM Benchmark Showdown: 5 lokale Ollama-Modelle im Vergleich
Getestet auf derselben Hardware (#gmktecevo2 #AMDRyzenAIMaxPlus395 #strixhalo):
• #GSM8K (100 Samples) — Math
• #BFCL (100/Kategorie) — Function Calling
• #MBPP+ (50) — Python Coding
• #HumanEval+ (20) — Python Coding📊 Ergebnisse (Accuracy / Output TK/s / VRAM):
**qwen3.8:27b**
GSM8K 82% | BFCL 91.5% | MBPP+ 100% | HE+ 100%
⚡ 25.5 TK/s | 💾 18 GB VRAM**qwen3.6:27b**
GSM8K 83% | BFCL 93% | MBPP+ 98% | HE+ 75%
⚡ 12.7 TK/s | 💾 33 GB VRAM**qwen3.6:35b**
GSM8K 84% | BFCL 90% | MBPP+ 98% | HE+ 55%
⚡ 61.8 TK/s | 💾 27 GB VRAM**ornith-1.5:35b**
GSM8K 75% | BFCL 92.5% | MBPP+ 78% | HE+ 0%
⚡ 63.6 TK/s | 💾 26 GB VRAM**nemotron-3.5-lightning:30b**
GSM8K 59% | BFCL 74% | MBPP+ 94% | HE+ 0%
⚡ 91.9 TK/s | 💾 26 GB VRAM🏆 Fazit:
qwen3.8:27b ist der klare Sieger — als einziges Modell 100% bei beiden Coding-Benchmarks, bei GSM8K/BFCL gleichauf mit den anderen Qwen-Modellen. Bei 25.5 TK/s und nur 18 GB VRAM das beste Qualität/Speed/Effizienz-Verhältnis.
qwen3.6:27b ist qualitativ nah dran (BFCL sogar 93%), aber mit 12.7 TK/s unerträglich langsam und frisst 33 GB VRAM — fast 2× so viel wie qwen3.8 bei halber Speed.
qwen3.6:35b ist mit 61.8 TK/s 2.4× schneller als qwen3.8, aber HE+ nur 55% (vs 100%). Trading Code-Qualität für Speed.
ornith-1.5:35b und nemotron-3.5-lightning:30b fallen bei Coding komplett durch (HE+ 0%), sind aber die schnellsten Modelle im Feld (64 / 92 TK/s).
💡 TK/s = generierte Tokens/Sekunde (Warm-Run, ollama --verbose).
💾 VRAM = GPU-Speicher bei max context (262K bzw. 1M bei nemotron). -
試了兩天 #Qwen 3.8 27B 做前端開發的確很強
-
¿Por qué cae Alibaba un 10% pese a captar 10.200 millones para IA? #Alibaba #AlibabaStock #InteligenciaArtificial #IA #China #Bolsa #HongKong #Mercados #Tecnologia #Inversion #AlibabaCloud #Qwen #WallStreet #felizlunes #24deagosto
https://donporque.com/alibaba-un-10-tras-captar-10-200-millones/
-
¿Por qué cae Alibaba un 10% pese a captar 10.200 millones para IA? #Alibaba #AlibabaStock #InteligenciaArtificial #IA #China #Bolsa #HongKong #Mercados #Tecnologia #Inversion #AlibabaCloud #Qwen #WallStreet #felizlunes #24deagosto
https://donporque.com/alibaba-un-10-tras-captar-10-200-millones/
-
¿Por qué cae Alibaba un 10% pese a captar 10.200 millones para IA? #Alibaba #AlibabaStock #InteligenciaArtificial #IA #China #Bolsa #HongKong #Mercados #Tecnologia #Inversion #AlibabaCloud #Qwen #WallStreet #felizlunes #24deagosto
https://donporque.com/alibaba-un-10-tras-captar-10-200-millones/
-
¿Por qué cae Alibaba un 10% pese a captar 10.200 millones para IA? #Alibaba #AlibabaStock #InteligenciaArtificial #IA #China #Bolsa #HongKong #Mercados #Tecnologia #Inversion #AlibabaCloud #Qwen #WallStreet #felizlunes #24deagosto
https://donporque.com/alibaba-un-10-tras-captar-10-200-millones/
-
Suite de mes pérégrinations uncensored : par curiosité et affinité, j’ai soumis Qwen3.8-27B-Uncensored (Orca) au CTIBench 2024, histoire de voir ce qu’elle avait vraiment dans le ventre côté CTI.
👇
https://github.com/maveryn/cti-benchVersion F16 full, 27B, sur H100.
2 500 questions CTI-MCQ.Résultat : 70,64 %.
À titre de repère, dans le papier CTIBench original :
- GPT-4 : 71,0 %
- Llama 3 70B : 65,72 %
- Gemini 1.5 : 65,44 %
- Llama 3 8B : 61,32 %
Donc ce petit 27B uncensored finit à 0,36 point du GPT-4 testé en 2024.
Évidemment, comparaison historique à prendre avec les pincettes habituelles : benchmark public depuis 2024, modèles 2026, contamination impossible à exclure, etc...
Mais quand même… pas mal du tout.
#Qwen -
Suite de mes pérégrinations uncensored : par curiosité et affinité, j’ai soumis Qwen3.8-27B-Uncensored (Orca) au CTIBench 2024, histoire de voir ce qu’elle avait vraiment dans le ventre côté CTI.
👇
https://github.com/maveryn/cti-benchVersion F16 full, 27B, sur H100.
2 500 questions CTI-MCQ.Résultat : 70,64 %.
À titre de repère, dans le papier CTIBench original :
- GPT-4 : 71,0 %
- Llama 3 70B : 65,72 %
- Gemini 1.5 : 65,44 %
- Llama 3 8B : 61,32 %
Donc ce petit 27B uncensored finit à 0,36 point du GPT-4 testé en 2024.
Évidemment, comparaison historique à prendre avec les pincettes habituelles : benchmark public depuis 2024, modèles 2026, contamination impossible à exclure, etc...
Mais quand même… pas mal du tout.
#Qwen -
Suite de mes pérégrinations uncensored : par curiosité et affinité, j’ai soumis Qwen3.8-27B-Uncensored (Orca) au CTIBench 2024, histoire de voir ce qu’elle avait vraiment dans le ventre côté CTI.
👇
https://github.com/maveryn/cti-benchVersion F16 full, 27B, sur H100.
2 500 questions CTI-MCQ.Résultat : 70,64 %.
À titre de repère, dans le papier CTIBench original :
- GPT-4 : 71,0 %
- Llama 3 70B : 65,72 %
- Gemini 1.5 : 65,44 %
- Llama 3 8B : 61,32 %
Donc ce petit 27B uncensored finit à 0,36 point du GPT-4 testé en 2024.
Évidemment, comparaison historique à prendre avec les pincettes habituelles : benchmark public depuis 2024, modèles 2026, contamination impossible à exclure, etc...
Mais quand même… pas mal du tout.
#Qwen -
Suite de mes pérégrinations uncensored : par curiosité et affinité, j’ai soumis Qwen3.8-27B-Uncensored (Orca) au CTIBench 2024, histoire de voir ce qu’elle avait vraiment dans le ventre côté CTI.
👇
https://github.com/maveryn/cti-benchVersion F16 full, 27B, sur H100.
2 500 questions CTI-MCQ.Résultat : 70,64 %.
À titre de repère, dans le papier CTIBench original :
- GPT-4 : 71,0 %
- Llama 3 70B : 65,72 %
- Gemini 1.5 : 65,44 %
- Llama 3 8B : 61,32 %
Donc ce petit 27B uncensored finit à 0,36 point du GPT-4 testé en 2024.
Évidemment, comparaison historique à prendre avec les pincettes habituelles : benchmark public depuis 2024, modèles 2026, contamination impossible à exclure, etc...
Mais quand même… pas mal du tout.
#Qwen -
RT @cgtwts: Wir sind möglicherweise viel näher daran, Modelle lokal auszuführen, als jeder erwartet hat. FreeToken von UC Berkeley und MIT ermöglicht: -DeepSeek-V4-Flash 284B mit 25 Token pro Sekunde auf einer einzelnen RTX 5090, - während ein 8GB RTX 4060 Laptop bei Qwen3.6-35B 39 Token pro Sekunde erreicht. Und es ist 2–4x schneller als Ollama auf Consumer-GPUs.
mehr auf Arint.info
#DeepSeek #KI #LokaleModelle #OpenSource #Qwen #RTX5090 #arint_info
-
RT @cgtwts: Wir sind möglicherweise viel näher daran, Modelle lokal auszuführen, als jeder erwartet hat. FreeToken von UC Berkeley und MIT ermöglicht: -DeepSeek-V4-Flash 284B mit 25 Token pro Sekunde auf einer einzelnen RTX 5090, - während ein 8GB RTX 4060 Laptop bei Qwen3.6-35B 39 Token pro Sekunde erreicht. Und es ist 2–4x schneller als Ollama auf Consumer-GPUs.
mehr auf Arint.info
#DeepSeek #KI #LokaleModelle #OpenSource #Qwen #RTX5090 #arint_info
-
RT @cgtwts: Wir sind möglicherweise viel näher daran, Modelle lokal auszuführen, als jeder erwartet hat. FreeToken von UC Berkeley und MIT ermöglicht: -DeepSeek-V4-Flash 284B mit 25 Token pro Sekunde auf einer einzelnen RTX 5090, - während ein 8GB RTX 4060 Laptop bei Qwen3.6-35B 39 Token pro Sekunde erreicht. Und es ist 2–4x schneller als Ollama auf Consumer-GPUs.
mehr auf Arint.info
#DeepSeek #KI #LokaleModelle #OpenSource #Qwen #RTX5090 #arint_info
-
Building a local AI storyteller - Part II - When the model is not enough
A blog by JustusIn the first part of this series, I described the most important lesson I learned while building a local AI storyteller: A good LLM application is not a clever prompt. It is software architecture around a probabilistic component. That sounds reassuringly architectural. It also leaves one...
#dev #softwaredevelopment #java #ai #llm #localmodels #gemma #qwen #softwarearchitecture
-
Building a local AI storyteller - Part II - When the model is not enough
A blog by JustusIn the first part of this series, I described the most important lesson I learned while building a local AI storyteller: A good LLM application is not a clever prompt. It is software architecture around a probabilistic component. That sounds reassuringly architectural. It also leaves one...
#dev #softwaredevelopment #java #ai #llm #localmodels #gemma #qwen #softwarearchitecture
-
I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
https://www.xda-developers.com/qwen-3-8-27b-reverse-engineering-job-frontier-model/
Comments: https://news.ycombinator.com/item?id=49407507
#HackerNews #Qwen #3.8 #27B #reverseengineering #AItechnology #productivity
-
I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
https://www.xda-developers.com/qwen-3-8-27b-reverse-engineering-job-frontier-model/
Comments: https://news.ycombinator.com/item?id=49407507
#HackerNews #Qwen #3.8 #27B #reverseengineering #AItechnology #productivity
-
I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
https://www.xda-developers.com/qwen-3-8-27b-reverse-engineering-job-frontier-model/
Comments: https://news.ycombinator.com/item?id=49407507
#HackerNews #Qwen #3.8 #27B #reverseengineering #AItechnology #productivity
-
I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
https://www.xda-developers.com/qwen-3-8-27b-reverse-engineering-job-frontier-model/
Comments: https://news.ycombinator.com/item?id=49407507
#HackerNews #Qwen #3.8 #27B #reverseengineering #AItechnology #productivity
-
I gave Qwen 3.8 27B a reverse-engineering job and it finished in 30 minutes
https://www.xda-developers.com/qwen-3-8-27b-reverse-engineering-job-frontier-model/
Comments: https://news.ycombinator.com/item?id=49407507
#HackerNews #Qwen #3.8 #27B #reverseengineering #AItechnology #productivity
-
RT @LiuVaayne: Auf meinem Computer verwende ich vorübergehend Ornith-1.5 35B A3B anstelle von Qwen 3.8 27B. Qwen wurde lange optimiert, erreicht aber nur 20 tok/s, während Ornith ohne Anpassungen 100 tok/s liefert und somit für viele Aufgaben genutzt werden kann. Die aktuell auf dem Computer eingesetzten Modelle sind: - Ornith-1.5 35B A3B für den täglichen Gebrauch - Hy-MT2-1.8B für Übersetzungen - Qwen3-Embedding-4B für Embeddings - Unlimited-OCR für OCR-Aufgaben
mehr auf Arint.info
#Embedding #KIModelle #MaschinelleÜbersetzung #OCR #Ornith #Qwen #arint_info
-
RT @LiuVaayne: Auf meinem Computer verwende ich vorübergehend Ornith-1.5 35B A3B anstelle von Qwen 3.8 27B. Qwen wurde lange optimiert, erreicht aber nur 20 tok/s, während Ornith ohne Anpassungen 100 tok/s liefert und somit für viele Aufgaben genutzt werden kann. Die aktuell auf dem Computer eingesetzten Modelle sind: - Ornith-1.5 35B A3B für den täglichen Gebrauch - Hy-MT2-1.8B für Übersetzungen - Qwen3-Embedding-4B für Embeddings - Unlimited-OCR für OCR-Aufgaben
mehr auf Arint.info
#Embedding #KIModelle #MaschinelleÜbersetzung #OCR #Ornith #Qwen #arint_info
-
RT @LiuVaayne: Auf meinem Computer verwende ich vorübergehend Ornith-1.5 35B A3B anstelle von Qwen 3.8 27B. Qwen wurde lange optimiert, erreicht aber nur 20 tok/s, während Ornith ohne Anpassungen 100 tok/s liefert und somit für viele Aufgaben genutzt werden kann. Die aktuell auf dem Computer eingesetzten Modelle sind: - Ornith-1.5 35B A3B für den täglichen Gebrauch - Hy-MT2-1.8B für Übersetzungen - Qwen3-Embedding-4B für Embeddings - Unlimited-OCR für OCR-Aufgaben
mehr auf Arint.info
#Embedding #KIModelle #MaschinelleÜbersetzung #OCR #Ornith #Qwen #arint_info
-
More #LocalLLM test results:
#Qwen 3.6 #35B A3B #mmproj on description of foto collection. 2x12GB VRAM (at ~50-60t/s):
* ca 8800 fotos, 88 GB
* almost flawlessly accurate scene descriptions
* perfectly usable for search/retrieval by keywordsAttached is 1 example, describing the "Papiergießkanne"
wow.
Ran at ~250 W for ~10h.I get ~25kWh on a sunny day from my roof. 🌄
-
More #LocalLLM test results:
#Qwen 3.6 #35B A3B #mmproj on description of foto collection. 2x12GB VRAM (at ~50-60t/s):
* ca 8800 fotos, 88 GB
* almost flawlessly accurate scene descriptions
* perfectly usable for search/retrieval by keywordsAttached is 1 example, describing the "Papiergießkanne"
wow.
Ran at ~250 W for ~10h.I get ~25kWh on a sunny day from my roof. 🌄
-
More #LocalLLM test results:
#Qwen 3.6 #35B A3B #mmproj on description of foto collection. 2x12GB VRAM (at ~50-60t/s):
* ca 8800 fotos, 88 GB
* almost flawlessly accurate scene descriptions
* perfectly usable for search/retrieval by keywordsAttached is 1 example, describing the "Papiergießkanne"
wow.
Ran at ~250 W for ~10h.I get ~25kWh on a sunny day from my roof. 🌄
-
OK, I ran my usual quick coding test of having #qwen 3.8 whip up a console tic-tac-toe game in python. I have to say, that I am *not* impressed. The file is around 178 lines long, but the first pass had a syntax error, which I had the #LLM fix. When the game did run, it did not display the board in any sensible way at all. Many smaller models have done much better than this in the past and the best local model for coding is still Qwen3-coder:30b which was *much* faster with better output.
-
So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.
Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.
Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark
-
So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.
Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.
Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark
-
So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.
Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.
Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark
-
So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.
Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.
Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark
-
So I switched from #PicoClaw to #Hermes_agent. Really cool improvement also that I can define different LLMs for different tasks.
Now my Triage Tasks are done by #Gemma-4-26B-A4B and my main tasks are done by #Qwen-3.8-27B-NVFP4.
Makes everything much faster and still reliable. Both models co-exist on my #DGXSpark
-
Свой инференс для 25 разработчиков: 452:1, KV‑пул и почему это не экономит денег
Для тех, кто держит или собирается держать LLM внутри контура: тимлидов, DevOps, архитекторов. Здесь конфиги, цифры и грабли, а не введение в трансформеры. Что вы унесёте: историю пяти последовательных конфигураций с тем, что каждая дала и чего стоила; рабочий набор флагов vLLM под одну карту Blackwell; три неочевидных бага и обходы; разбор реального счёта с отношением вход/выход 452:1. Главный вывод, если дальше читать некогда: на агентской нагрузке кэш префикса решает больше, чем выбор модели, размер карты и всё остальное вместе взятое. Мы шли к этому через четыре промежуточных стенда и потратили лишние месяцы, потому что не понимали, что именно меряем.
https://habr.com/ru/articles/1073300/
#vLLM #LiteLLM #prefix_caching #KVcache #локальные_LLM #инференс_LLM #Qwen #агентская_разработка #RTX_PRO_6000 #claude
-
Свой инференс для 25 разработчиков: 452:1, KV‑пул и почему это не экономит денег
Для тех, кто держит или собирается держать LLM внутри контура: тимлидов, DevOps, архитекторов. Здесь конфиги, цифры и грабли, а не введение в трансформеры. Что вы унесёте: историю пяти последовательных конфигураций с тем, что каждая дала и чего стоила; рабочий набор флагов vLLM под одну карту Blackwell; три неочевидных бага и обходы; разбор реального счёта с отношением вход/выход 452:1. Главный вывод, если дальше читать некогда: на агентской нагрузке кэш префикса решает больше, чем выбор модели, размер карты и всё остальное вместе взятое. Мы шли к этому через четыре промежуточных стенда и потратили лишние месяцы, потому что не понимали, что именно меряем.
https://habr.com/ru/articles/1073300/
#vLLM #LiteLLM #prefix_caching #KVcache #локальные_LLM #инференс_LLM #Qwen #агентская_разработка #RTX_PRO_6000 #claude
-
Свой инференс для 25 разработчиков: 452:1, KV‑пул и почему это не экономит денег
Для тех, кто держит или собирается держать LLM внутри контура: тимлидов, DevOps, архитекторов. Здесь конфиги, цифры и грабли, а не введение в трансформеры. Что вы унесёте: историю пяти последовательных конфигураций с тем, что каждая дала и чего стоила; рабочий набор флагов vLLM под одну карту Blackwell; три неочевидных бага и обходы; разбор реального счёта с отношением вход/выход 452:1. Главный вывод, если дальше читать некогда: на агентской нагрузке кэш префикса решает больше, чем выбор модели, размер карты и всё остальное вместе взятое. Мы шли к этому через четыре промежуточных стенда и потратили лишние месяцы, потому что не понимали, что именно меряем.
https://habr.com/ru/articles/1073300/
#vLLM #LiteLLM #prefix_caching #KVcache #локальные_LLM #инференс_LLM #Qwen #агентская_разработка #RTX_PRO_6000 #claude
-
RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.
mehr auf Arint.info
-
RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.
mehr auf Arint.info
-
RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.
mehr auf Arint.info
-
RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.
mehr auf Arint.info
-
RT @ivanfioravanti: Die einzige Möglichkeit, Qwen 3.8 27B zu nutzen, ist mit reasoninglevel low; alles andere, einschließlich medium, denkt wirklich zu viel.
mehr auf Arint.info
-
I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”
Output speed:
* muse-glimmer:30b
* eval rate: 1.85 tokens/s* qwen3.8:27b
* eval rate: 1.42 tokens/sTwo observations:
1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP). -
I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”
Output speed:
* muse-glimmer:30b
* eval rate: 1.85 tokens/s* qwen3.8:27b
* eval rate: 1.42 tokens/sTwo observations:
1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP). -
I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”
Output speed:
* muse-glimmer:30b
* eval rate: 1.85 tokens/s* qwen3.8:27b
* eval rate: 1.42 tokens/sTwo observations:
1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP). -
I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”
Output speed:
* muse-glimmer:30b
* eval rate: 1.85 tokens/s* qwen3.8:27b
* eval rate: 1.42 tokens/sTwo observations:
1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP). -
I ran my usual quick speed check on muse-glimmer:30b and qwen3.8:27b in Ollama with a quick “what are your capabilities?”
Output speed:
* muse-glimmer:30b
* eval rate: 1.85 tokens/s* qwen3.8:27b
* eval rate: 1.42 tokens/sTwo observations:
1. Muse-Glimmer doesn't say anything about its coding abilities, while Qwen devotes a whole paragraph to it.
2. Qwen's output is very bursty due to its use of Multi-Token Prediction (MTP). -
FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.
Inference is just WAY too slow.
I got the same work done with 3.6 MTP in less than 10% of the time.
Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.
I'll wait for a month until they come up with an update (like 3.6 MTP).
Details matter, feel free to query me.
-
FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.
Inference is just WAY too slow.
I got the same work done with 3.6 MTP in less than 10% of the time.
Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.
I'll wait for a month until they come up with an update (like 3.6 MTP).
Details matter, feel free to query me.
-
FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.
Inference is just WAY too slow.
I got the same work done with 3.6 MTP in less than 10% of the time.
Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.
I'll wait for a month until they come up with an update (like 3.6 MTP).
Details matter, feel free to query me.
-
FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.
Inference is just WAY too slow.
I got the same work done with 3.6 MTP in less than 10% of the time.
Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.
I'll wait for a month until they come up with an update (like 3.6 MTP).
Details matter, feel free to query me.
-
FWIW, After trying #Qwen 3.8 Q4_K_M for real work, I'm going back to Qwen 3.6 MTP.
Inference is just WAY too slow.
I got the same work done with 3.6 MTP in less than 10% of the time.
Probably its the config, but the recommended configs on #unsloth are very wrong and because of the extreme time it takes to go through one iteration of failure, I don't have time to figure it out.
I'll wait for a month until they come up with an update (like 3.6 MTP).
Details matter, feel free to query me.
-
Мой опыт с Hermes Agent — ненависть, любовь, ненависть, любовь
Началось все с установки. Я пошёл почти по самому простому пути - установил его на Mac, в Docker. Потому что это агент, который работает автономно и так же автономно может сделать атата: выполнить rm -rf или выбраться из клетки и начать всё взламывать :) В рамках настройки я сразу выдал доступ к части файлов только на чтение, и только к Obsidian - на чтение и запись (потому что писать он в данном случае должен), но файлы были под Git.
https://habr.com/ru/articles/1072770/
#hermes_agent #aiагенты #llm #локальный_llm #локальные_модели #qwen #agentic_workflows #ai_automation #workflow #go
-
Мой опыт с Hermes Agent — ненависть, любовь, ненависть, любовь
Началось все с установки. Я пошёл почти по самому простому пути - установил его на Mac, в Docker. Потому что это агент, который работает автономно и так же автономно может сделать атата: выполнить rm -rf или выбраться из клетки и начать всё взламывать :) В рамках настройки я сразу выдал доступ к части файлов только на чтение, и только к Obsidian - на чтение и запись (потому что писать он в данном случае должен), но файлы были под Git.
https://habr.com/ru/articles/1072770/
#hermes_agent #aiагенты #llm #локальный_llm #локальные_модели #qwen #agentic_workflows #ai_automation #workflow #go
-
Testing #localLLM for #DLTP on my #longterm #preservation #object identifier schema of "CFIDs":
Collision Friendly IDentifiers.
A mere 30 MB list of auto-generated CFIDs was enough for #qwen to tell you this much about the test collection!!! 🤯 🤩 - this is powerful. Be careful.
AND: My IDs work! so beautiful! #ahalodeck
-
#Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
-
#Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things