#qwen3 — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #qwen3, aggregated by home.social.
-
Спекулятивное декодирование: от чего на самом деле зависит ускорение
Спекулятивное декодирование обещает ускорение до 6.5 раза: маленькая черновая модель набрасывает несколько токенов вперёд, большая проверяет их за один свой проход, а текст на выходе получается тот же самый. Я воспроизвёл EAGLE-3 на бесплатной Kaggle T4 и получил разброс от 2.6 раза на коде до полного отсутствия выигрыша вне домена, на котором обучалась черновая голова, — при одной и той же реализации, на одной карте, без единой подкрутки. Разбираю, какая величина этим разбросом управляет, почему глубина дерева черновиков способна увести метод в минус и какие пять вопросов стоит задавать к любой цифре вида «в N раз быстрее».
https://habr.com/ru/articles/1082452/
#спекулятивное_декодирование #EAGLE3 #llm #инференс #ускорение_инференса #qwen3 #длина_принятия
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
https://narilabs.com/blog/nari-labs-leads-coval-voice-ai-benchmarks/
Comments: https://news.ycombinator.com/item?id=49699267
#HackerNews #NariQwen3 #TTS #Qwen3 #ASR #VoiceAI #LowLatency
-
RT @TeksEdge: Jemand hat das größte Problem von Qwen3.8-27B behoben: das Überdenken! Ich muss das unbedingt ausprobieren!! ÜBERDENKEN.
mehr auf Arint.info
#AI #DeepLearning #Innovation #MachineLearning #Qwen3 #TechNews #arint_info
-
RT @TeksEdge: Jemand hat das größte Problem von Qwen3.8-27B behoben: das Überdenken! Ich muss das unbedingt ausprobieren!! ÜBERDENKEN.
mehr auf Arint.info
#AI #DeepLearning #Innovation #MachineLearning #Qwen3 #TechNews #arint_info
-
eben krasser vergleich #qwen3.8max vs #glm53
ich arbeite mit qwen3.8max
will ein ollama pull ausführen:Sorry, da hat sich eine Schleife eingeschlichen — notify=true als Boolean wird gerade nicht akzeptiert, und ich hab stur wiederholt statt umzuschalten. Mein Fehler.
Mach es kurz selbst im Terminal:
ollama pull tobestyledintro/nex-n2.5-mini:q4_k_m
Sag mir Bescheid wenn der Pull durch ist, dann starte ich sofort den Smoke-Test (Tool-Calling + BFCL-Stichprobe).er will also das ich das selbst mache weil er es nicht hin bekommt
okay - ich pull manuell - stolper danach auf #hugging_face noch über anderes llm was ich testen will:
qwen:
Ich hab wieder das gleiche notify-Problem wie vorhin. Mach den Download bitte selbst kurz im Terminal:huggingface-cli download agentionai/Signal-3.8-27B-GGUF \
--include "AP-Q4_K_XL/*.gguf" --local-dir ~/models/signal-3.8
Wenn der Download fertig ist, sag Bescheid — dann erstelle ich die Modelfile und importiere es in Ollama. Der Nex-Pull läuft ja parallel bei dir.ich sag ihm: löse das problem mal bei dir: probiert stur das gleiche aus
⚠️ Tool guardrail halted terminal: identical_call_streak_halt
ich schalte um auf GLM:
Ja, da ist was faul — und zwar an mir: Ich hab immer wieder den gleichen notify-Parameter mitgeschickt, den die Toolbox scheinbar nicht frisst, statt ihn einfach wegzulassen. Die Lösung ist trivial: notify komplett weglassen und den Prozess selbst pollen. Mach ich jetzt:
💻 terminal
hf download agentionai/Signal-3.8-27B...
Läuft. Jetzt poll ich den Download:
⚙️ process_manage: "wait proc_1212bbc8eeb 60s"
💻 terminaldu -sh ~/models/signal-3.8/ >/dev/nu...
⚙️ process_manage: "wait proc_1212bbc8eeb 600s"Download läuft — 443 MB von ~16,35 GiB nach ~1 Minute. Das wird noch ein paar Minuten dauern. Ich melde Lars kurz den Status und warte dann weiter.
WARUM kommt Qwen nicht auf die Idee??? glm ist deshalb einfach das beste =)