#qwen3 — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #qwen3, aggregated by home.social.
-
Спекулятивное декодирование: от чего на самом деле зависит ускорение
Спекулятивное декодирование обещает ускорение до 6.5 раза: маленькая черновая модель набрасывает несколько токенов вперёд, большая проверяет их за один свой проход, а текст на выходе получается тот же самый. Я воспроизвёл EAGLE-3 на бесплатной Kaggle T4 и получил разброс от 2.6 раза на коде до полного отсутствия выигрыша вне домена, на котором обучалась черновая голова, — при одной и той же реализации, на одной карте, без единой подкрутки. Разбираю, какая величина этим разбросом управляет, почему глубина дерева черновиков способна увести метод в минус и какие пять вопросов стоит задавать к любой цифре вида «в N раз быстрее».
https://habr.com/ru/articles/1082452/
#спекулятивное_декодирование #EAGLE3 #llm #инференс #ускорение_инференса #qwen3 #длина_принятия
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
RT @TeksEdge: Jemand hat das größte Problem von Qwen3.8-27B behoben: das Überdenken! Ich muss das unbedingt ausprobieren!! ÜBERDENKEN.
mehr auf Arint.info
#AI #DeepLearning #Innovation #MachineLearning #Qwen3 #TechNews #arint_info
-
RT @TeksEdge: Jemand hat das größte Problem von Qwen3.8-27B behoben: das Überdenken! Ich muss das unbedingt ausprobieren!! ÜBERDENKEN.
mehr auf Arint.info
#AI #DeepLearning #Innovation #MachineLearning #Qwen3 #TechNews #arint_info
-
RT @0xBakeer: Qwen3.8-Flash-Next auf einer einzelnen DGX Spark: bis zu 97 Tokens pro Sekunde.
mehr auf Arint.info
#AI #DGXSpark #MachineLearning #NLP #OpenSource #Qwen3 #arint_info
-
🚀🎉 Wow, it's truly groundbreaking news that #Qwen3.8 can churn out a whole 37 tokens per second on a #GPU you'd need a #loan to buy! 💸 Apparently, this "experiment" showed that throwing every kitchen sink at a model doesn't exactly equal genius-level #AI. 🙄
https://piszczek.pl/blog/qwen38-27b-256k-50-tps-24gb-gpu #breakthrough #AIexperiment #technews #HackerNews #ngated -
Just ran Qwen3-4B locally on LM Studio and asked "who are you?"
It said: "I am Gemini, a large language model developed by Google."
Sir. You are a 2.5GB GGUF file sitting in my RAM. You have never seen the internet. Google has no idea you exist.
Local AI is wild. 🏠🤖
#LocalAI #LMStudio #Qwen3 #LLM #AIHallucination #selfhosted