home.social

#qwen3 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #qwen3, aggregated by home.social.

  1. 🚀 Wow, "Qwen 3.8" is now hot on the heels of "GPT-5.5 Pro" in the prestigious field of... *reasoning prefills* 🤯. Because, let's face it, who doesn't love a good prefilled reasoning session from the vast wilderness of #GitHub gists? 😜
    gist.github.com/wsxiaoys/e0286 #Qwen3.8 #GPT5.5 #reasoningprefills #gists #AItechnology #HackerNews #ngated

  2. 🚀 Wow, "Qwen 3.8" is now hot on the heels of "GPT-5.5 Pro" in the prestigious field of... *reasoning prefills* 🤯. Because, let's face it, who doesn't love a good prefilled reasoning session from the vast wilderness of #GitHub gists? 😜
    gist.github.com/wsxiaoys/e0286 #Qwen3.8 #GPT5.5 #reasoningprefills #gists #AItechnology #HackerNews #ngated

  3. 🚀 Wow, "Qwen 3.8" is now hot on the heels of "GPT-5.5 Pro" in the prestigious field of... *reasoning prefills* 🤯. Because, let's face it, who doesn't love a good prefilled reasoning session from the vast wilderness of #GitHub gists? 😜
    gist.github.com/wsxiaoys/e0286 #Qwen3.8 #GPT5.5 #reasoningprefills #gists #AItechnology #HackerNews #ngated

  4. 🚀 Wow, "Qwen 3.8" is now hot on the heels of "GPT-5.5 Pro" in the prestigious field of... *reasoning prefills* 🤯. Because, let's face it, who doesn't love a good prefilled reasoning session from the vast wilderness of #GitHub gists? 😜
    gist.github.com/wsxiaoys/e0286 #Qwen3.8 #GPT5.5 #reasoningprefills #gists #AItechnology #HackerNews #ngated

  5. 🚀 Wow, "Qwen 3.8" is now hot on the heels of "GPT-5.5 Pro" in the prestigious field of... *reasoning prefills* 🤯. Because, let's face it, who doesn't love a good prefilled reasoning session from the vast wilderness of #GitHub gists? 😜
    gist.github.com/wsxiaoys/e0286 #Qwen3.8 #GPT5.5 #reasoningprefills #gists #AItechnology #HackerNews #ngated

  6. RT @TeksEdge: ❔Erinnert ihr euch an das Gigabyte AORUS RTX-5060TI eGPU für 699$? Nun ... 🤯Ein Reddit-Nutzer (bygiolasagna) hat es geschafft, das neue Qwen3.8-27B vollständig in eine einzelne RTX 5060 Ti mit 16GB zu pressen. Und es läuft mit fast 48 Token pro Sekunde. Es handelt sich um ein dichtes 27B-Modell. Die Konfiguration: 🔥 RTX 5060 Ti — 16GB 🧠 Qwen3.8-27B dense 🗜️ ~14.60 GiB benutzerdefiniertes IQ4XS GGUF 🦙 llama.cpp CUDA ⚡ Vollständiges GPU-Offloading 🚀 Flash Attention + CUDA Graphs 📚 32K Kontextfenster 💾 Q4 KV-Cache 🧩 MTP-2 Performance: 🐢 Ohne MTP → 25,7 tok/s ⚡ MTP-1 → 40,0 tok/s 🚀 MTP-2 → 47,4–47,6 tok/s Selbst nach einem Prefill von ~30K Token: → 45,7 tok/s Das entspricht einer Steigerung von etwa 85% durch MTP. 👀 Es ist bemerkenswert, dass ein 27B dichtes multimodales/agentic-Modell vollständig auf einer 16GB-Verbraucher-GPU der 700$-Klasse resident ist und mit wirklich interaktiven Geschwindigkeiten läuft.

    mehr auf Arint.info

    #AI #Hardware #LLM #MachineLearning #Qwen3 #RTX5060Ti #arint_info

    https://x.com/TeksEdge/status/2089155254118170941#m

  7. Welcome to the magical world of Qwen 3.8 Max Preview, where AI models are now cheaper than your morning latte, and you can seamlessly integrate your toolchain (whatever that means). 🥳🎉 Just remember, you'll need a PhD in Qwen Settings to actually get started, but hey, deep thinking is overrated anyway. 🤖🧠
    qwencloud.com/pricing/token-pl #Qwen3.8MaxPreview #AIModels #AffordableTech #ToolchainIntegration #DeepThinking #HackerNews #ngated

  8. Сжатие декодерных эмбеддеров: как ужать 8B до продакшена без потери recall

    Декодерный эмбеддер 7–8B дает качество, но платит за него памятью, latency и деньгами. Разбираем все оси сжатия - int8, int4, binary + rescoring, PQ, MRL-усечение - на реальных замерах recall@10: где деградация мягкая, а где обрыв. С воспроизводимым кодом и Colab-ноутбуком под Qwen3

    habr.com/ru/articles/1054930/

    #сжатие_эмбеддингов #квантизация #эмбеддинги #embeddings #RAG #Qdrant #Qwen3 #binary_quantization #Matryoshka #retrieval

  9. Making the most out of a small LLM

    Yesterday i finally built my own #AI #server. I had a spare #Nvidia RTX 2070 with 8GB of #VRAM laying around and wanted to do this for a long time.

    The problem is that most #LLMs need a lot of VRAM and i don't want to buy another #GPU just to host my own AI. Then i came across #gemma3 and #qwen3. Both of these are amazing #quantized models with stunning reasoning given that they need so less resources.

    I chose huihui_ai/qwen3-abliterated:14b since it supports #deepthinking, #toolcalling and is pretty unrestricted. After some testing i noticed that the 8b model performs even better than the 14b variant with drastically better performance. I can't make out any quality loss there to be honest. The 14b model sneaked in chinese characters into the response very often. The 8b model on the other hand doesn't.

    Now i've got a very fast model with amazing reasoning (even in German) and tool calling support. The only thing left to improve is knowledge. #Firecrawl is a great tool for #webscraping and as soon as i implemented websearching, the setup was complete. At least i thought it was.

    I want to make the most out of this LLM and therefore my next step is to implement a basic #webserver that exposes the same #API #endpoints as #ollama so that everywhere ollama is supported, i can point it to my python script instead. This way it feels like the model is way more capable than it actually is. I can use these advanced features everywhere without being bound to it's actual knowledge.

    To improve this setup even more i will likely switch to a #mixture_of_experts architecture soon. This project is a lot of fun and i can't wait to integrate it into my homelab.

    #homelab #selfhosting #privacy #ai #llm #largelanguagemodels #coding #developement