#unsloth — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #unsloth, aggregated by home.social.
-
RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: https://github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: https://unsloth.ai/docs/basics/amd
mehr auf Arint.info
#AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info
-
RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: https://github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: https://unsloth.ai/docs/basics/amd
mehr auf Arint.info
#AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info
-
RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: https://github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: https://unsloth.ai/docs/basics/amd
mehr auf Arint.info
#AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info
-
RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: https://github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: https://unsloth.ai/docs/basics/amd
mehr auf Arint.info
#AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info
-
RT @UnslothAI: Wir stellen Unsloth für AMD vor 🚀 Sie können jetzt LLMs auf Ihrer AMD-Hardware trainieren und ausführen • Wir haben mit AMD zusammengearbeitet, um Ihnen das Training und die Ausführung von über 500 Modellen auf AMD-GPUs zu ermöglichen • Funktioniert unter Windows, WSL, Linux • Trainieren Sie Qwen, Gemma auf 3GB VRAM GitHub: https://github.com/unslothai/unsloth Funktioniert auf Radeon, Instinct, Ryzen und Data-Center-GPUs mit bis zu 2× schnellerer Leistung bei 70% weniger VRAM und ohne Genauigkeitsverlust dank unserer benutzerdefinierten Triton-Kernel und mathematischen Algorithmen. Wir unterstützen auch optimierte ROCm-Builds für GGUF & Safetensors-Inference. Unsloth ist eine Open-Source-Lokalschnittstelle für schnelleres LLM-Training und Inference, mit Tool-Call-Healing, Code-Ausführung, sicherer Websuche, Remote-APIs und HTTPS-Bereitstellung. Verbinden Sie lokale Modelle mit Claude Code, Codex-Agenten und führen Sie die neuesten Kimi-, GLM-, DeepSeek-, Qwen3.6- und Gemma-4-Modelle aus. 🔗Blog + Anleitung: https://unsloth.ai/docs/basics/amd
mehr auf Arint.info
#AI #AMD #LLM #MachineLearning #OpenSource #Unsloth #arint_info
-
-
RT @UnslothAI: We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU. Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B hits 13K tok/s (B200). Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference. Blog: https://unsloth.ai/docs/basics/nvfp4 Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
-
RT @UnslothAI: We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU. Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B hits 13K tok/s (B200). Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference. Blog: https://unsloth.ai/docs/basics/nvfp4 Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
-
RT @UnslothAI: We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU. Gemma-4-12B NVFP4 works on 11GB VRAM. 26B-A4B hits 13K tok/s (B200). Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference. Blog: https://unsloth.ai/docs/basics/nvfp4 Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
RT @UnslothAI: We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4
mehr auf Arint.info
#agent #huggingface #Qwen36 #qwen36 #Qwen3627 #unsloth #arint_info
-
Как я обучил русский RAG‑сплиттер, который режет документы по индексам, а не по тексту
TL;DR. Из интереса обучил собственный русский RAG‑сплиттер — захотелось проверить, можно ли сделать context‑aware‑нарезку русских документов лучше готовых чанкеров. Я взял идею датской context-aware-splitter , пересобрал её под русский на базе T-lite-it-2.1 и изменил главное: модель возвращает индексы границ, а не переписанный текст. Хост потом режет оригинал по этим индексам. У index‑output оказалось три практических плюса:
https://habr.com/ru/articles/1055628/
#rag #чанкинг #дистилляция #lora #unsloth #токенизация #llamacpp #gguf #vulkan #amd
-
Как я обучил русский RAG‑сплиттер, который режет документы по индексам, а не по тексту
TL;DR. Из интереса обучил собственный русский RAG‑сплиттер — захотелось проверить, можно ли сделать context‑aware‑нарезку русских документов лучше готовых чанкеров. Я взял идею датской context-aware-splitter , пересобрал её под русский на базе T-lite-it-2.1 и изменил главное: модель возвращает индексы границ, а не переписанный текст. Хост потом режет оригинал по этим индексам. У index‑output оказалось три практических плюса:
https://habr.com/ru/articles/1055628/
#rag #чанкинг #дистилляция #lora #unsloth #токенизация #llamacpp #gguf #vulkan #amd
-
Как я обучил русский RAG‑сплиттер, который режет документы по индексам, а не по тексту
TL;DR. Из интереса обучил собственный русский RAG‑сплиттер — захотелось проверить, можно ли сделать context‑aware‑нарезку русских документов лучше готовых чанкеров. Я взял идею датской context-aware-splitter , пересобрал её под русский на базе T-lite-it-2.1 и изменил главное: модель возвращает индексы границ, а не переписанный текст. Хост потом режет оригинал по этим индексам. У index‑output оказалось три практических плюса:
https://habr.com/ru/articles/1055628/
#rag #чанкинг #дистилляция #lora #unsloth #токенизация #llamacpp #gguf #vulkan #amd
-
Do you feel the itch to run a state of the art #LLM at home? With the MIT licensed GLM-5.2 and a bit of spare RAM you now can! Thank you #unsloth for those articles!
https://unsloth.ai/docs/models/glm-5.2 -
Same week, small update: Run LLMs Locally
Multi-Token-Prediction (MTP) for Gemma-4-E4B and Gemma-4-26B from Unsloth. After 50% from QAT, this brings another 25-90% improvement in token generation speed.
The OpenCode config slide received a small update to reduce prompt sizes with "rtk" and "opencode-tool-search", reducing default prompt size by 60 percent.
Also added logging all prompts to the parameter list.https://codeberg.org/thbley/talks/raw/branch/main/Run_LLMs_Locally_2026_ThomasBley.pdf
-
Same week, small update: Run LLMs Locally
Multi-Token-Prediction (MTP) for Gemma-4-E4B and Gemma-4-26B from Unsloth. After 50% from QAT, this brings another 25-90% improvement in token generation speed.
The OpenCode config slide received a small update to reduce prompt sizes with "rtk" and "opencode-tool-search", reducing default prompt size by 60 percent.
Also added logging all prompts to the parameter list.https://codeberg.org/thbley/talks/raw/branch/main/Run_LLMs_Locally_2026_ThomasBley.pdf
-
How Unsloth and Nvidia made LLM training 25% faster on consumer GPUs
https://unsloth.ai/blog/nvidia-collab
#HackerNews #Unsloth #Nvidia #LLMtraining #ConsumerGPUs #AItechnology
-
How Unsloth and Nvidia made LLM training 25% faster on consumer GPUs
https://unsloth.ai/blog/nvidia-collab
#HackerNews #Unsloth #Nvidia #LLMtraining #ConsumerGPUs #AItechnology
-
How Unsloth and Nvidia made LLM training 25% faster on consumer GPUs
https://unsloth.ai/blog/nvidia-collab
#HackerNews #Unsloth #Nvidia #LLMtraining #ConsumerGPUs #AItechnology
-
How Unsloth and Nvidia made LLM training 25% faster on consumer GPUs
https://unsloth.ai/blog/nvidia-collab
#HackerNews #Unsloth #Nvidia #LLMtraining #ConsumerGPUs #AItechnology
-
How Unsloth and Nvidia made LLM training 25% faster on consumer GPUs
https://unsloth.ai/blog/nvidia-collab
#HackerNews #Unsloth #Nvidia #LLMtraining #ConsumerGPUs #AItechnology
-
[Перевод] Локальный запуск GLM-5.1
Перевод подготовил автор канала Друг Опенсурса , приятного прочтения, заранее благодарю за подписку В этой статье мы подробно разберем процесс развертывания GLM-5.1 с использованием llama.cpp и форматов GGUF. Узнаем о системных требованиях, сборке и настройках, оптимизации и практическом применении.
https://habr.com/ru/articles/1022242/
#glm51 #llm #Llamacpp #Unsloth #GGUF #Локальный_запуск #tool_calling #Zai #искусственный_интеллект
-
[Перевод] Локальный запуск GLM-5.1
Перевод подготовил автор канала Друг Опенсурса , приятного прочтения, заранее благодарю за подписку В этой статье мы подробно разберем процесс развертывания GLM-5.1 с использованием llama.cpp и форматов GGUF. Узнаем о системных требованиях, сборке и настройках, оптимизации и практическом применении.
https://habr.com/ru/articles/1022242/
#glm51 #llm #Llamacpp #Unsloth #GGUF #Локальный_запуск #tool_calling #Zai #искусственный_интеллект
-
[Перевод] Локальный запуск GLM-5.1
Перевод подготовил автор канала Друг Опенсурса , приятного прочтения, заранее благодарю за подписку В этой статье мы подробно разберем процесс развертывания GLM-5.1 с использованием llama.cpp и форматов GGUF. Узнаем о системных требованиях, сборке и настройках, оптимизации и практическом применении.
https://habr.com/ru/articles/1022242/
#glm51 #llm #Llamacpp #Unsloth #GGUF #Локальный_запуск #tool_calling #Zai #искусственный_интеллект
-
Unsloth's "guide" to fine-tuning #Qwen3.5 is the digital equivalent of watching paint dry, but with more #acronyms and chevrons. 🤦♂️🚀 If you enjoy reading a glorified list of links and buzzwords, you'll be in heaven—otherwise, pray for a faster MoE escape plan! 🏃♀️💨
https://unsloth.ai/docs/models/qwen3.5/fine-tune #Unsloth #guide #techboredom #digitalescape #HackerNews #ngated -
Unsloth's "guide" to fine-tuning #Qwen3.5 is the digital equivalent of watching paint dry, but with more #acronyms and chevrons. 🤦♂️🚀 If you enjoy reading a glorified list of links and buzzwords, you'll be in heaven—otherwise, pray for a faster MoE escape plan! 🏃♀️💨
https://unsloth.ai/docs/models/qwen3.5/fine-tune #Unsloth #guide #techboredom #digitalescape #HackerNews #ngated -
Fine-tuning Qwen-8B под проприетарный синтаксис (CADINP) на одной RTX 3090: опыт инженера-конструктора
Возможно ли на одной домашней видеокарте (RTX 3090) создать AI-ассистента, который знает узкоспециализированный инженерный язык лучше, чем GPT-4? Я инженер-конструктор, и мне надоело писать рутинный код для SOFiSTiK руками. Поэтому я решил дообучить (fine-tune) модель Qwen 3 (8B) с дистилляцией логики DeepSeek под свои задачи. В статье подробный технический разбор: — Как собрать датасет с логикой Chain of Thought (CoT). — Как бороться с Out of Memory в 24 ГБ VRAM на Windows + WSL. — Рабочие конфиги Unsloth, параметры обучения и итоговая GGUF модель. Раскрыть
https://habr.com/ru/articles/987240/
#LLM #finetuning #локальные_нейросети #RTX_3090 #Unsloth #Qwen #DeepSeek #GGUF #SOFiSTiK #CADINP
-
Fine-tuning Qwen-8B под проприетарный синтаксис (CADINP) на одной RTX 3090: опыт инженера-конструктора
Возможно ли на одной домашней видеокарте (RTX 3090) создать AI-ассистента, который знает узкоспециализированный инженерный язык лучше, чем GPT-4? Я инженер-конструктор, и мне надоело писать рутинный код для SOFiSTiK руками. Поэтому я решил дообучить (fine-tune) модель Qwen 3 (8B) с дистилляцией логики DeepSeek под свои задачи. В статье подробный технический разбор: — Как собрать датасет с логикой Chain of Thought (CoT). — Как бороться с Out of Memory в 24 ГБ VRAM на Windows + WSL. — Рабочие конфиги Unsloth, параметры обучения и итоговая GGUF модель. Раскрыть
https://habr.com/ru/articles/987240/
#LLM #finetuning #локальные_нейросети #RTX_3090 #Unsloth #Qwen #DeepSeek #GGUF #SOFiSTiK #CADINP
-
Fine-tuning Qwen-8B под проприетарный синтаксис (CADINP) на одной RTX 3090: опыт инженера-конструктора
Возможно ли на одной домашней видеокарте (RTX 3090) создать AI-ассистента, который знает узкоспециализированный инженерный язык лучше, чем GPT-4? Я инженер-конструктор, и мне надоело писать рутинный код для SOFiSTiK руками. Поэтому я решил дообучить (fine-tune) модель Qwen 3 (8B) с дистилляцией логики DeepSeek под свои задачи. В статье подробный технический разбор: — Как собрать датасет с логикой Chain of Thought (CoT). — Как бороться с Out of Memory в 24 ГБ VRAM на Windows + WSL. — Рабочие конфиги Unsloth, параметры обучения и итоговая GGUF модель. Раскрыть
https://habr.com/ru/articles/987240/
#LLM #finetuning #локальные_нейросети #RTX_3090 #Unsloth #Qwen #DeepSeek #GGUF #SOFiSTiK #CADINP
-
Fine-tuning Qwen-8B под проприетарный синтаксис (CADINP) на одной RTX 3090: опыт инженера-конструктора Возможно ли на одной ...
#LLM #fine-tuning #локальные #нейросети #RTX #3090 #Unsloth #Qwen #DeepSeek #GGUF #SOFiSTiK
Origin | Interest | Match -
MiniMax-M2 đã có mặt trên Hugging Face nhờ Unsloth! Bản guff có thể sớm ra mắt, mở ra khả năng sử dụng local dễ dàng hơn. 🚀
#LocalLLaMA #AI #MachineLearning #Unsloth #MiniMaxM2 #trituenhantao #hocmayhttps://www.reddit.com/r/LocalLLaMA/comments/1oibaz2/waiting_for_an_unsloth_guff_for_minimaxm2/
-
I am testing the capabilities of some small #LLM 's on #LMStudio today. These 7 to 12 B models are much stronger than I thought. Some of them run pretty fast, but some larger models are burning my #rtx4060 #GPU. I think I will settle with #IBM #Granite 3.3 which is a 8B model but was further trained by #unsloth to 9B. Granite 3.3 came out in April this year. In the long run, I will need a 20 to 40B model. But then I most likely need an rtx 5090 machine with 64G VRAM to run them. #AI #AIs
-
Train your own R1 reasoning model with Unsloth.
"We've enhanced the entire GRPO process, making it use 80% less VRAM than Hugging Face + FA2. This allows you to reproduce R1-Zero's "aha moment" on just 7GB of VRAM using Qwen2.5 (1.5B)"
#ai #reasoning #unsloth #opensource #locally #grpo
https://unsloth.ai/blog/r1-reasoning -
Train your own R1 reasoning model with Unsloth.
"We've enhanced the entire GRPO process, making it use 80% less VRAM than Hugging Face + FA2. This allows you to reproduce R1-Zero's "aha moment" on just 7GB of VRAM using Qwen2.5 (1.5B)"
#ai #reasoning #unsloth #opensource #locally #grpo
https://unsloth.ai/blog/r1-reasoning