home.social

#qwen3 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #qwen3, aggregated by home.social.

  1. 📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B

    🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B

  2. 📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B

    🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B

  3. 📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B

    🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B

  4. 📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B

    🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B

  5. New slides: Run LLMs Locally

    Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
    Added testing with curl for TypeSafe System One API compatible endpoints.
    Added text to speech with Qwen-3-TTS using llama.cpp.
    Added Nemotron 3.5 Lightning.

    github.com/thomas-0816/talks/b

    #ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev

  6. New slides: Run LLMs Locally

    Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
    Added testing with curl for TypeSafe System One API compatible endpoints.
    Added text to speech with Qwen-3-TTS using llama.cpp.
    Added Nemotron 3.5 Lightning.

    github.com/thomas-0816/talks/b

    #ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev

  7. New slides: Run LLMs Locally

    Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
    Added testing with curl for TypeSafe System One API compatible endpoints.
    Added text to speech with Qwen-3-TTS using llama.cpp.
    Added Nemotron 3.5 Lightning.

    github.com/thomas-0816/talks/b

    #ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev

  8. New slides: Run LLMs Locally

    Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
    Added testing with curl for TypeSafe System One API compatible endpoints.
    Added text to speech with Qwen-3-TTS using llama.cpp.
    Added Nemotron 3.5 Lightning.

    github.com/thomas-0816/talks/b

    #ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev

  9. RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.

    mehr auf Arint.info

    #AI #DDR3 #GPU #Hardware #Qwen3 #Strata #arint_info

    https://x.com/coldniko/status/2105657680303948063

  10. RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.

    mehr auf Arint.info

    #AI #DDR3 #GPU #Hardware #Qwen3 #Strata #arint_info

    https://x.com/coldniko/status/2105657680303948063

  11. RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.

    mehr auf Arint.info

    #AI #DDR3 #GPU #Hardware #Qwen3 #Strata #arint_info

    https://x.com/coldniko/status/2105657680303948063

  12. RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.

    mehr auf Arint.info

    #AI #DDR3 #GPU #Hardware #Qwen3 #Strata #arint_info

    https://x.com/coldniko/status/2105657680303948063