#qwen3 — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #qwen3, aggregated by home.social.
-
Small update: Run LLMs Locally
Added Maple-Preview 20B-A1B using llama.cpp (6 GB, fastest 20B model).
Added Qwen 3.8 27B configuration for coding agents.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #clef
-
Small update: Run LLMs Locally
Added Maple-Preview 20B-A1B using llama.cpp (6 GB, fastest 20B model).
Added Qwen 3.8 27B configuration for coding agents.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #clef
-
Small update: Run LLMs Locally
Added Maple-Preview 20B-A1B using llama.cpp (6 GB, fastest 20B model).
Added Qwen 3.8 27B configuration for coding agents.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #clef
-
Small update: Run LLMs Locally
Added Maple-Preview 20B-A1B using llama.cpp (6 GB, fastest 20B model).
Added Qwen 3.8 27B configuration for coding agents.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #clef
-
Celebrating #DigitalIndependenceDay with my latest shift: #AI used to mean poor quality, high energy use, and dependency on a few companies, so I avoided it. Self-hosted open-weight models like #Qwen3.8 now offer sovereignty, cost efficiency, and quality that suffices for agentic workflows. #OpenWebUI brings a feature-rich frontend, #LiteLLM gateways requests across providers, and #OpenCode is a provider-agnostic coding agent.
-
Celebrating #DigitalIndependenceDay with my latest shift: #AI used to mean poor quality, high energy use, and dependency on a few companies, so I avoided it. Self-hosted open-weight models like #Qwen3.8 now offer sovereignty, cost efficiency, and quality that suffices for agentic workflows. #OpenWebUI brings a feature-rich frontend, #LiteLLM gateways requests across providers, and #OpenCode is a provider-agnostic coding agent.
-
Celebrating #DigitalIndependenceDay with my latest shift: #AI used to mean poor quality, high energy use, and dependency on a few companies, so I avoided it. Self-hosted open-weight models like #Qwen3.8 now offer sovereignty, cost efficiency, and quality that suffices for agentic workflows. #OpenWebUI brings a feature-rich frontend, #LiteLLM gateways requests across providers, and #OpenCode is a provider-agnostic coding agent.
-
Celebrating #DigitalIndependenceDay with my latest shift: #AI used to mean poor quality, high energy use, and dependency on a few companies, so I avoided it. Self-hosted open-weight models like #Qwen3.8 now offer sovereignty, cost efficiency, and quality that suffices for agentic workflows. #OpenWebUI brings a feature-rich frontend, #LiteLLM gateways requests across providers, and #OpenCode is a provider-agnostic coding agent.
-
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
https://github.com/Niko1221/Strata
Comments: https://news.ycombinator.com/item?id=49953495
#HackerNews #Qwen3.8 #Flash #RTX4090 #ConsumerHardware #AIperformance #100T/s
-
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
https://github.com/Niko1221/Strata
Comments: https://news.ycombinator.com/item?id=49953495
#HackerNews #Qwen3.8 #Flash #RTX4090 #ConsumerHardware #AIperformance #100T/s
-
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
https://github.com/Niko1221/Strata
Comments: https://news.ycombinator.com/item?id=49953495
#HackerNews #Qwen3.8 #Flash #RTX4090 #ConsumerHardware #AIperformance #100T/s
-
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
https://github.com/Niko1221/Strata
Comments: https://news.ycombinator.com/item?id=49953495
#HackerNews #Qwen3.8 #Flash #RTX4090 #ConsumerHardware #AIperformance #100T/s
-
📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B
🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B
-
📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B
🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B
-
📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B
🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B
-
📏 262,144 tokens of native context, validated up to 1,048,576. 40 of 50 layers use 512-token sliding-window attention, every 5th layer attends to everything. RULER at 1M tokens: 63.2 vs 57.5 for #Qwen3.5 35B-A3B
🇩🇪 Reasons in German on German prompts, with 4 reasoning effort levels (none, low, medium, high). AIME 2025: 87.5 in German, 96.9 in English. Trained with the Merlin-Arthur protocol to abstain: it says "I don't know" 44% of the time instead of hallucinating, vs 11.1% for Qwen3.5 35B-A3B
-
Small update: Run LLMs Locally
Added decision model Clef-Flash using llama.cpp.
https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
More about Clef from Cloudflare: https://blog.cloudflare.com/clef-decision-models/
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev #clef
-
Small update: Run LLMs Locally
Added decision model Clef-Flash using llama.cpp.
https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
More about Clef from Cloudflare: https://blog.cloudflare.com/clef-decision-models/
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev #clef
-
Small update: Run LLMs Locally
Added decision model Clef-Flash using llama.cpp.
https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
More about Clef from Cloudflare: https://blog.cloudflare.com/clef-decision-models/
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev #clef
-
Small update: Run LLMs Locally
Added decision model Clef-Flash using llama.cpp.
https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
More about Clef from Cloudflare: https://blog.cloudflare.com/clef-decision-models/
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev #clef
-
New slides: Run LLMs Locally
Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
Added testing with curl for TypeSafe System One API compatible endpoints.
Added text to speech with Qwen-3-TTS using llama.cpp.
Added Nemotron 3.5 Lightning.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev
-
New slides: Run LLMs Locally
Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
Added testing with curl for TypeSafe System One API compatible endpoints.
Added text to speech with Qwen-3-TTS using llama.cpp.
Added Nemotron 3.5 Lightning.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev
-
New slides: Run LLMs Locally
Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
Added testing with curl for TypeSafe System One API compatible endpoints.
Added text to speech with Qwen-3-TTS using llama.cpp.
Added Nemotron 3.5 Lightning.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev
-
New slides: Run LLMs Locally
Added decision models Laya, lev, Kev, OpenJev using llama.cpp.
Added testing with curl for TypeSafe System One API compatible endpoints.
Added text to speech with Qwen-3-TTS using llama.cpp.
Added Nemotron 3.5 Lightning.https://github.com/thomas-0816/talks/blob/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #stablediffusion #qwen3 #qwen3tts #nemotron #glm #localai #opencode #systemone #laya #lev #kev #openjev
-
RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.
mehr auf Arint.info
-
RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.
mehr auf Arint.info
-
RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.
mehr auf Arint.info
-
RT @coldniko: Viele Menschen verstehen Strata immer noch nicht vollständig. Du kannst jetzt eine alte, günstige DDR3-Maschine mit 128 GB oder 64 GB RAM kaufen, eine 8 GB+ GPU anschließen und Qwen3.8-Flash-Next mit über 70 Tokens pro Sekunde betreiben. Denkt daran: Mehr Kerne/Threads in der CPU bedeuten schnellere Leistung.
mehr auf Arint.info