home.social

#unsloth — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #unsloth, aggregated by home.social.

  1. ah geil! #nvidia hat ein eigenes repo für #opensuse 😎

    jetzt hab ich endlich cuda 13.2 auf 13.3 updaten können um endlich mal #unsloth qwen3.6 mit #llamacpp laufen zu lassen

    #KI #homelab #linux #qwen

  2. ah geil! #nvidia hat ein eigenes repo für #opensuse 😎

    jetzt hab ich endlich cuda 13.2 auf 13.3 updaten können um endlich mal #unsloth qwen3.6 mit #llamacpp laufen zu lassen

    #KI #homelab #linux #qwen

  3. ah geil! #nvidia hat ein eigenes repo für #opensuse 😎

    jetzt hab ich endlich cuda 13.2 auf 13.3 updaten können um endlich mal #unsloth qwen3.6 mit #llamacpp laufen zu lassen

    #KI #homelab #linux #qwen

  4. ah geil! #nvidia hat ein eigenes repo für #opensuse 😎

    jetzt hab ich endlich cuda 13.2 auf 13.3 updaten können um endlich mal #unsloth qwen3.6 mit #llamacpp laufen zu lassen

    #KI #homelab #linux #qwen

  5. ah geil! #nvidia hat ein eigenes repo für #opensuse 😎

    jetzt hab ich endlich cuda 13.2 auf 13.3 updaten können um endlich mal #unsloth qwen3.6 mit #llamacpp laufen zu lassen

    #KI #homelab #linux #qwen

  6. RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: unsloth.ai/docs/basics/amd

    mehr auf Arint.info

    #AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info

    https://x.com/UnslothAI/status/2079207457788952944#m

  7. RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: unsloth.ai/docs/basics/amd

    mehr auf Arint.info

    #AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info

    https://x.com/UnslothAI/status/2079207457788952944#m

  8. RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: unsloth.ai/docs/basics/amd

    mehr auf Arint.info

    #AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info

    https://x.com/UnslothAI/status/2079207457788952944#m

  9. RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: unsloth.ai/docs/basics/amd

    mehr auf Arint.info

    #AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info

    https://x.com/UnslothAI/status/2079207457788952944#m

  10. RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: unsloth.ai/docs/basics/amd

    mehr auf Arint.info

    #AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info

    https://x.com/UnslothAI/status/2079207457788952944#m

  11. 8G VRAM的話我可能還是會選擇用 #Gemma4 26b 因為 #Unsloth 的Q4速度比它快又支持MTP,可能等llamacpp更新了對 Bonsai 支持更完善再測試,持續觀望

  12. #Bonsai 27b Ternary 2bit的感覺上與 #Unsloth 的Q2好像差不多,我16G VRAM的機器跑Q2已經不錯,所以就先測試它1bit的因為8G能跑

  13. #Unsloth 的 IQ2 K XL Qwen3.6 27b 性能出乎意料地高,而且能夠在16G VRAM裏跑

  14. RT @UnslothAI: We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU. Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B hits 13K tok/s (B200). Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference. Blog: unsloth.ai/docs/basics/nvfp4 Gemma NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #huggingface #unsloth #Unsloth #arint_info

    https://x.com/UnslothAI/status/2077070408948584617#m

  15. RT @UnslothAI: We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU. Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B hits 13K tok/s (B200). Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference. Blog: unsloth.ai/docs/basics/nvfp4 Gemma NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #huggingface #unsloth #Unsloth #arint_info

    https://x.com/UnslothAI/status/2077070408948584617#m

  16. RT @UnslothAI: We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU. Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B hits 13K tok/s (B200). Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference. Blog: unsloth.ai/docs/basics/nvfp4 Gemma NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #huggingface #unsloth #Unsloth #arint_info

    https://x.com/UnslothAI/status/2077070408948584617#m

  17. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  18. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  19. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  20. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  21. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  22. RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: unsloth.ai/docs/models/qwen3.6 Qwen3.6 NVFP4: huggingface.co/collections/uns

    mehr auf Arint.info

    #agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info

    https://x.com/UnslothAI/status/2075566124687892597#m

  23. Как я обучил русский RAG‑сплиттер, который режет документы по индексам, а не по тексту

    TL;DR. Из интереса обучил собственный русский RAG‑сплиттер — захотелось проверить, можно ли сделать context‑aware‑нарезку русских документов лучше готовых чанкеров. Я взял идею датской context-aware-splitter , пересобрал её под русский на базе T-lite-it-2.1 и изменил главное: модель возвращает индексы границ, а не переписанный текст. Хост потом режет оригинал по этим индексам. У index‑output оказалось три практических плюса:

    habr.com/ru/articles/1055628/

    #rag #чанкинг #дистилляция #lora #unsloth #токенизация #llamacpp #gguf #vulkan #amd

  24. Как я обучил русский RAG‑сплиттер, который режет документы по индексам, а не по тексту

    TL;DR. Из интереса обучил собственный русский RAG‑сплиттер — захотелось проверить, можно ли сделать context‑aware‑нарезку русских документов лучше готовых чанкеров. Я взял идею датской context-aware-splitter , пересобрал её под русский на базе T-lite-it-2.1 и изменил главное: модель возвращает индексы границ, а не переписанный текст. Хост потом режет оригинал по этим индексам. У index‑output оказалось три практических плюса:

    habr.com/ru/articles/1055628/

    #rag #чанкинг #дистилляция #lora #unsloth #токенизация #llamacpp #gguf #vulkan #amd

  25. Как я обучил русский RAG‑сплиттер, который режет документы по индексам, а не по тексту

    TL;DR. Из интереса обучил собственный русский RAG‑сплиттер — захотелось проверить, можно ли сделать context‑aware‑нарезку русских документов лучше готовых чанкеров. Я взял идею датской context-aware-splitter , пересобрал её под русский на базе T-lite-it-2.1 и изменил главное: модель возвращает индексы границ, а не переписанный текст. Хост потом режет оригинал по этим индексам. У index‑output оказалось три практических плюса:

    habr.com/ru/articles/1055628/

    #rag #чанкинг #дистилляция #lora #unsloth #токенизация #llamacpp #gguf #vulkan #amd

  26. Do you feel the itch to run a state of the art #LLM at home? With the MIT licensed GLM-5.2 and a bit of spare RAM you now can! Thank you #unsloth for those articles!
    unsloth.ai/docs/models/glm-5.2

  27. Same week, small update: Run LLMs Locally

    Multi-Token-Prediction (MTP) for Gemma-4-E4B and Gemma-4-26B from Unsloth. After 50% from QAT, this brings another 25-90% improvement in token generation speed.

    The OpenCode config slide received a small update to reduce prompt sizes with "rtk" and "opencode-tool-search", reducing default prompt size by 60 percent.
    Also added logging all prompts to the parameter list.

    codeberg.org/thbley/talks/raw/

    #ai #llm #llamacpp #localai #gemma4 #opencode #mtp #unsloth

  28. Same week, small update: Run LLMs Locally

    Multi-Token-Prediction (MTP) for Gemma-4-E4B and Gemma-4-26B from Unsloth. After 50% from QAT, this brings another 25-90% improvement in token generation speed.

    The OpenCode config slide received a small update to reduce prompt sizes with "rtk" and "opencode-tool-search", reducing default prompt size by 60 percent.
    Also added logging all prompts to the parameter list.

    codeberg.org/thbley/talks/raw/

    #ai #llm #llamacpp #localai #gemma4 #opencode #mtp #unsloth

  29. Unsloth Gemma 4 QAT: Some deep-in-the-weeds details about boiling down an LLM to a small size you can run on a single desktop computer (or phone)
    unsloth.ai/docs/models/gemma-4
    #unsloth #google #quant #llm #ai #+

  30. [Перевод] Локальный запуск GLM-5.1

    Перевод подготовил автор канала Друг Опенсурса , приятного прочтения, заранее благодарю за подписку В этой статье мы подробно разберем процесс развертывания GLM-5.1 с использованием llama.cpp и форматов GGUF. Узнаем о системных требованиях, сборке и настройках, оптимизации и практическом применении.

    habr.com/ru/articles/1022242/

    #glm51 #llm #Llamacpp #Unsloth #GGUF #Локальный_запуск #tool_calling #Zai #искусственный_интеллект

  31. [Перевод] Локальный запуск GLM-5.1

    Перевод подготовил автор канала Друг Опенсурса , приятного прочтения, заранее благодарю за подписку В этой статье мы подробно разберем процесс развертывания GLM-5.1 с использованием llama.cpp и форматов GGUF. Узнаем о системных требованиях, сборке и настройках, оптимизации и практическом применении.

    habr.com/ru/articles/1022242/

    #glm51 #llm #Llamacpp #Unsloth #GGUF #Локальный_запуск #tool_calling #Zai #искусственный_интеллект

  32. [Перевод] Локальный запуск GLM-5.1

    Перевод подготовил автор канала Друг Опенсурса , приятного прочтения, заранее благодарю за подписку В этой статье мы подробно разберем процесс развертывания GLM-5.1 с использованием llama.cpp и форматов GGUF. Узнаем о системных требованиях, сборке и настройках, оптимизации и практическом применении.

    habr.com/ru/articles/1022242/

    #glm51 #llm #Llamacpp #Unsloth #GGUF #Локальный_запуск #tool_calling #Zai #искусственный_интеллект

  33. Unsloth's "guide" to fine-tuning #Qwen3.5 is the digital equivalent of watching paint dry, but with more #acronyms and chevrons. 🤦‍♂️🚀 If you enjoy reading a glorified list of links and buzzwords, you'll be in heaven—otherwise, pray for a faster MoE escape plan! 🏃‍♀️💨
    unsloth.ai/docs/models/qwen3.5 #Unsloth #guide #techboredom #digitalescape #HackerNews #ngated

  34. Unsloth's "guide" to fine-tuning #Qwen3.5 is the digital equivalent of watching paint dry, but with more #acronyms and chevrons. 🤦‍♂️🚀 If you enjoy reading a glorified list of links and buzzwords, you'll be in heaven—otherwise, pray for a faster MoE escape plan! 🏃‍♀️💨
    unsloth.ai/docs/models/qwen3.5 #Unsloth #guide #techboredom #digitalescape #HackerNews #ngated

  35. Fine-tuning Qwen-8B под проприетарный синтаксис (CADINP) на одной RTX 3090: опыт инженера-конструктора

    Возможно ли на одной домашней видеокарте (RTX 3090) создать AI-ассистента, который знает узкоспециализированный инженерный язык лучше, чем GPT-4? Я инженер-конструктор, и мне надоело писать рутинный код для SOFiSTiK руками. Поэтому я решил дообучить (fine-tune) модель Qwen 3 (8B) с дистилляцией логики DeepSeek под свои задачи. В статье подробный технический разбор: — Как собрать датасет с логикой Chain of Thought (CoT). — Как бороться с Out of Memory в 24 ГБ VRAM на Windows + WSL. — Рабочие конфиги Unsloth, параметры обучения и итоговая GGUF модель. Раскрыть

    habr.com/ru/articles/987240/

    #LLM #finetuning #локальные_нейросети #RTX_3090 #Unsloth #Qwen #DeepSeek #GGUF #SOFiSTiK #CADINP

  36. Fine-tuning Qwen-8B под проприетарный синтаксис (CADINP) на одной RTX 3090: опыт инженера-конструктора

    Возможно ли на одной домашней видеокарте (RTX 3090) создать AI-ассистента, который знает узкоспециализированный инженерный язык лучше, чем GPT-4? Я инженер-конструктор, и мне надоело писать рутинный код для SOFiSTiK руками. Поэтому я решил дообучить (fine-tune) модель Qwen 3 (8B) с дистилляцией логики DeepSeek под свои задачи. В статье подробный технический разбор: — Как собрать датасет с логикой Chain of Thought (CoT). — Как бороться с Out of Memory в 24 ГБ VRAM на Windows + WSL. — Рабочие конфиги Unsloth, параметры обучения и итоговая GGUF модель. Раскрыть

    habr.com/ru/articles/987240/

    #LLM #finetuning #локальные_нейросети #RTX_3090 #Unsloth #Qwen #DeepSeek #GGUF #SOFiSTiK #CADINP

  37. Fine-tuning Qwen-8B под проприетарный синтаксис (CADINP) на одной RTX 3090: опыт инженера-конструктора

    Возможно ли на одной домашней видеокарте (RTX 3090) создать AI-ассистента, который знает узкоспециализированный инженерный язык лучше, чем GPT-4? Я инженер-конструктор, и мне надоело писать рутинный код для SOFiSTiK руками. Поэтому я решил дообучить (fine-tune) модель Qwen 3 (8B) с дистилляцией логики DeepSeek под свои задачи. В статье подробный технический разбор: — Как собрать датасет с логикой Chain of Thought (CoT). — Как бороться с Out of Memory в 24 ГБ VRAM на Windows + WSL. — Рабочие конфиги Unsloth, параметры обучения и итоговая GGUF модель. Раскрыть

    habr.com/ru/articles/987240/

    #LLM #finetuning #локальные_нейросети #RTX_3090 #Unsloth #Qwen #DeepSeek #GGUF #SOFiSTiK #CADINP

  38. Fine-tuning Qwen-8B под проприетарный синтаксис (CADINP) на одной RTX 3090: опыт инженера-конструктора Возможно ли на одной ...

    #LLM #fine-tuning #локальные #нейросети #RTX #3090 #Unsloth #Qwen #DeepSeek #GGUF #SOFiSTiK

    Origin | Interest | Match
  39. I am testing the capabilities of some small #LLM 's on #LMStudio today. These 7 to 12 B models are much stronger than I thought. Some of them run pretty fast, but some larger models are burning my #rtx4060 #GPU. I think I will settle with #IBM #Granite 3.3 which is a 8B model but was further trained by #unsloth to 9B. Granite 3.3 came out in April this year. In the long run, I will need a 20 to 40B model. But then I most likely need an rtx 5090 machine with 64G VRAM to run them. #AI #AIs

  40. Train your own R1 reasoning model with Unsloth.
    "We've enhanced the entire GRPO process, making it use 80% less VRAM than Hugging Face + FA2. This allows you to reproduce R1-Zero's "aha moment" on just 7GB of VRAM using Qwen2.5 (1.5B)"
    #ai #reasoning #unsloth #opensource #locally #grpo
    unsloth.ai/blog/r1-reasoning

  41. Train your own R1 reasoning model with Unsloth.
    "We've enhanced the entire GRPO process, making it use 80% less VRAM than Hugging Face + FA2. This allows you to reproduce R1-Zero's "aha moment" on just 7GB of VRAM using Qwen2.5 (1.5B)"
    #ai #reasoning #unsloth #opensource #locally #grpo
    unsloth.ai/blog/r1-reasoning