home.social

#qwen3 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #qwen3, aggregated by home.social.

  1. RT @imnotchalk: Qwen3.8 27B ist jetzt im @cerebras Shared Tier verfügbar mit einer Geschwindigkeit von ~1500 Token pro Sekunde. 🚀

    mehr auf Arint.info

    #AI #Cerebras #LLM #MachineLearning #Qwen3 #TechNews #arint_info

    https://x.com/imnotchalk/status/2095637567979114654

  2. RT @imnotchalk: Qwen3.8 27B ist jetzt im @cerebras Shared Tier verfügbar mit einer Geschwindigkeit von ~1500 Token pro Sekunde. 🚀

    mehr auf Arint.info

    #AI #Cerebras #LLM #MachineLearning #Qwen3 #TechNews #arint_info

    https://x.com/imnotchalk/status/2095637567979114654

  3. RT @imnotchalk: Qwen3.8 27B ist jetzt im @cerebras Shared Tier verfügbar mit einer Geschwindigkeit von ~1500 Token pro Sekunde. 🚀

    mehr auf Arint.info

    #AI #Cerebras #LLM #MachineLearning #Qwen3 #TechNews #arint_info

    https://x.com/imnotchalk/status/2095637567979114654

  4. RT @imnotchalk: Qwen3.8 27B ist jetzt im @cerebras Shared Tier verfügbar mit einer Geschwindigkeit von ~1500 Token pro Sekunde. 🚀

    mehr auf Arint.info

    #AI #Cerebras #LLM #MachineLearning #Qwen3 #TechNews #arint_info

    https://x.com/imnotchalk/status/2095637567979114654

  5. RT @imnotchalk: Qwen3.8 27B ist jetzt im @cerebras Shared Tier verfügbar mit einer Geschwindigkeit von ~1500 Token pro Sekunde. 🚀

    mehr auf Arint.info

    #AI #Cerebras #LLM #MachineLearning #Qwen3 #TechNews #arint_info

    https://x.com/imnotchalk/status/2095637567979114654

  6. 🔍🎨 So, it seems Qwen 3.8 is the new Picasso of #AI, with its 27 billion parameters ready to dazzle—or confuse—you with a dizzying array of credits and models that sound like a fruit salad 🍌🍎. Who knew API testing involved so many Nano Bananas and Seedreams? At least now we know #ByteDance isn't just for TikTok dances; it’s also for pricey AI models. 💸💃
    imageat.com/models/qwen-3-8-27 #Art #Qwen3.8 #APItesting #TechTrends #HackerNews #ngated

  7. RT @ciruai: Das beste Qwen3.8 27B-Modell für AMD Strix Halo ist da. In Zusammenarbeit mit Pete Hopton und dem Kairic.ai-Team präsentiere ich: Qwen3.8-27B-IU4-KAIRIC-EDGE. Unserer Kenntnis nach ist dies die weltweit erste öffentliche Veröffentlichung eines 27B-LLM mit einer beschleunigten nativen IU4-Inferenzspur, speziell für AMD Strix Halo (gfx1151) entwickelt. Warum ist IU4 wichtig? Die meisten 4-Bit-Modelle speichern zwar Speicher, konvertieren Gewichte jedoch weiterhin in breitere Formate während der Ausführung. Unsere IU4-Spur hält Vier-Bit-Ganzzahldaten im aktiven Berechnungspfad und mappt sie direkt auf RDNA 3.5-Integer-Hardware. Das reduziert den Speichertraffic und verwandelt Quantisierung von einem Speicherformat in eine echte Hardware-Ausführungsstrategie. Der abgestimmte IU4-Operator erreichte 104,66 TOPS – 1,94× FP16 und 1,93× IU8. Dies eröffnet einen Weg zu schnellerer, effizienterer lokaler KI, ohne dabei die Modellqualität zu opfern. Ich kann es kaum erwarten zu sehen, was @Italianclownz und andere Zauberer damit anstellen. Angetrieben von Prompt Forge + Dual View, Kairic Edge-Beschleunigung, 262K Kontext: HumanEval-Metriken: • 47,73 tok/s Full-Suite TG • +85% gegenüber Unsloth Dynamic Q4 • +89% gegenüber Unsloth Dynamic Q6 • Schlug Unsloth Dynamic v3 q6 in Human Eval um 2 Fragen. Dies ist auch die erste Veröffentlichung, die von Kairic.ai angetrieben wird, einem neuen Unternehmen, das sich auf die gemeinsame Entwicklung von KI-Software rund um die Hardware konzentriert, auf der sie tatsächlich läuft. Modell, Benchmarks, Quelle, Build-Anleitung und Runner: huggingface.c…

    mehr auf Arint.info

    #AMDStrixHalo #KairicAI #LLM #LocalAI #Qwen3 #RDNA3 #arint_info

    https://x.com/ciruai/status/2091004843720659229

  8. RT @ciruai: Das beste Qwen3.8 27B-Modell für AMD Strix Halo ist da. In Zusammenarbeit mit Pete Hopton und dem Kairic.ai-Team präsentiere ich: Qwen3.8-27B-IU4-KAIRIC-EDGE. Unserer Kenntnis nach ist dies die weltweit erste öffentliche Veröffentlichung eines 27B-LLM mit einer beschleunigten nativen IU4-Inferenzspur, speziell für AMD Strix Halo (gfx1151) entwickelt. Warum ist IU4 wichtig? Die meisten 4-Bit-Modelle speichern zwar Speicher, wandeln Gewichte jedoch während der Ausführung weiterhin in breitere Formate um. Unsere IU4-Spur behält die Vier-Bit-Ganzzahldaten im aktiven Berechnungspfad und leitet sie direkt an die RDNA 3.5-Ganzzahlhardware weiter. Das reduziert den Speichertraffic und verwandelt Quantisierung von einem Speicherformat in eine echte Hardware-Ausführungsstrategie. Der abgestimmte IU4-Operator erreichte 104,66 TOPS – 1,94× FP16 und 1,93× IU8. Dies eröffnet einen Weg zu schnellerer, effizienterer lokaler KI, ohne dabei die Modellqualität zu opfern. Ich kann es kaum erwarten zu sehen, was @Italianclownz und andere Experten damit anstellen werden. Getrieben von Prompt Forge + Dual View, Kairic Edge-Beschleunigung, 262K Kontext: HumanEval-Metriken: • 47,73 tok/s Full-Suite TG • +85% gegenüber Unsloth Dynamic Q4 • +89% gegenüber Unsloth Dynamic Q6 • Schlug Unsloth Dynamic v3 q6 in HumanEval um 2 Fragen. Dies ist auch die erste Veröffentlichung, die von Kairic.ai unterstützt wird, einem neuen Unternehmen, das sich auf die gemeinsame Entwicklung von KI-Software rund um die Hardware konzentriert, auf der sie tatsächlich läuft. Modell, Benchmarks, Quelle, Build-Anleitung und Ru…

    mehr auf Arint.info

    #AMDStrixHalo #KairicAI #LLM #LocalAI #Qwen3 #RDNA3 #arint_info

    https://x.com/ciruai/status/2091004843720659229

  9. Update, more slides: Run LLMs Locally

    I added sandboxing of OpenCode and llama.cpp with nono and Landlock.
    And a new slide to describe jailbreaks with DeepInception.

    github.com/thomas-0816/talks/b

    #ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode

  10. New week, new slides: Run LLMs Locally

    I added virtualization of OpenCode with Matchlock and Firecracker microVMs,
    containerization of OpenCode and llama.cpp with Docker
    and a new slide for indirect prompt injection attacks.
    Matchlock is a great project for sandboxing, bringing the advantages of containers to virtual machines.

    github.com/thomas-0816/talks/b

    #ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode #firecracker #docker

  11. New week, new slides and small updates: Run LLMs Locally

    Added an example to create Mermaid diagrams in llama.cpp UI.
    Added QAT (Quantization-Aware Training) variants of Gemma 4 which are 50 percent faster in token generation with my local setup.
    Added definitions for Deterministic and Probabilistic results.

    github.com/thomas-0816/talks/b

    #ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode #mtp #webassembly #mellum2

  12. New week, beautiful new slides: Run LLMs Locally

    Now with Mellum2 from JetBrains!
    A very fast coding model, requires only 10 GB RAM.

    I also added LFM 2.5 from LiquidAI, updated translations with HY-MT2 from Tencent, added examples for wllama using re-ranking and structured output
    and added thinking_budget_tokens to the curl examples.

    github.com/thomas-0816/talks/b

    #ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode #mtp #webassembly #jetbrains #mellum2

  13. New week, more slides: Run LLMs Locally

    Now including wllama to run GGUF models inside your browser!

    wllama uses llama.cpp, WebAssembly and WebGPU, bringing a completely new experience of LLMs into the web.
    It has no 4 GB limitation and is faster than Transformers.js.

    I also added translations using the HY-MT model from Tencent.

    github.com/thomas-0816/talks/b

    #ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode #mtp #webassembly

  14. Erkenntnis nach Selbstversuch: Wenn acht Sekunden Sample ausreichen, um lokale KI einen aktuellen Text mit meiner Stimme vorlesen zu lassen - das ändert in Sachen Quellenglaubwürdigkeit absehbar alles.

    #qwen3 #selbstversuch #ai #journalism auch im #lokaljournalismus

  15. Have updated my local #LLM server from #Mistral 8B to #Qwen3 30B A3B. Still very fast. "Thinking" model so it tends to follow instructions and handle prompts with copious amounts of data better (#HomeAssistant reports).

    This is all done on CPU, too. 20 cores, but still. If you have #HomeLab equipment needing a use, I highly recommend setting up #oobabooga and slapping one of these bad bois in there so you can call it via API in your other projects.