#qwen3 — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #qwen3, aggregated by home.social.
-
📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B
🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B
-
📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B
🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B
-
📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B
🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B
-
📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B
🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B
-
Small update: Run LLMs Locally
Added decision model Clef-Flash using llama.cpp.
https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
More about Clef from Cloudflare: https://blog.cloudflare.com/clef-decision-models/
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev #clef
-
Small update: Run LLMs Locally
Added decision model Clef-Flash using llama.cpp.
https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
More about Clef from Cloudflare: https://blog.cloudflare.com/clef-decision-models/
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev #clef
-
Small update: Run LLMs Locally
Added decision model Clef-Flash using llama.cpp.
https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
More about Clef from Cloudflare: https://blog.cloudflare.com/clef-decision-models/
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev #clef
-
Small update: Run LLMs Locally
Added decision model Clef-Flash using llama.cpp.
https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
More about Clef from Cloudflare: https://blog.cloudflare.com/clef-decision-models/
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev #clef
-
New slides: Run LLMs Locally
Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
Added testing with curl for TypeSafe System One API compatible endpoints.
Added text to speech with Qwen-3-TTS using llama.cpp.
Added Nemotron 3.5 Lightning.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev
-
New slides: Run LLMs Locally
Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
Added testing with curl for TypeSafe System One API compatible endpoints.
Added text to speech with Qwen-3-TTS using llama.cpp.
Added Nemotron 3.5 Lightning.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev
-
New slides: Run LLMs Locally
Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
Added testing with curl for TypeSafe System One API compatible endpoints.
Added text to speech with Qwen-3-TTS using llama.cpp.
Added Nemotron 3.5 Lightning.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev
-
New slides: Run LLMs Locally
Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
Added testing with curl for TypeSafe System One API compatible endpoints.
Added text to speech with Qwen-3-TTS using llama.cpp.
Added Nemotron 3.5 Lightning.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev
-
RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.
mehr auf Arint.info
-
RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.
mehr auf Arint.info
-
RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.
mehr auf Arint.info
-
RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.
mehr auf Arint.info