#rocm — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #rocm, aggregated by home.social.
-
I got Deepseek V4 work in Framework Desktop, an AMD Strix Halo PC with 128G memory, inside Lemonade server. I briefly tested code snippet generation, and result was pretty good. Downsides are that it's rather slow and context is short. But coming from a mini-PC it's still *very* impressive. The little Framework PC keeps giving.
Speed was about 15 tps, but the speed stays constantly there - even when context gets longer. I remember running Kimi K2 a year ago at 3 tps!
I tested code generation by asking it to implement an Angular module for OAUTH login client. I refined it through few iterations to add e.g. hardening and configuration features. Code quality was very good. Finally i asked it to write it out as plan.md, restarted and asked to generate code from the plan. Regenerated code was nearly identical to original round.
The server had some instability after chat grew to about 35k long (total 15k tokens). Nothing crashes but client showed an error that stream ended prematurely. The server log showed it finished though. Overall this was good experience, with some concern about actual max content length.
Model was unsloth/DeepSeek-V4-Flash-0731-GGUF with UD-IQ2_M quant. The Lemonade server couldn't run it out-of-the-box, complaining about unknown "Deepseek" architecture.
- I upgraded llama.cpp to a nightly build:
lemonade config set llamacpp.rocm_bin=b10230
- Reloaded the llama.cpp backend from Lemonade UI.
- File/Add model.
- Set run parameters: --flash-attn on --reasoning on -np 1 --ctx-checkpoints 0
- Max context: 65kRAM usage was at 90GB, so there is still room for another model in parallel, or better quant. No crashes, even after several hours. Next step is try Hermes Studio with the model.
#homelab #AI #deepseek #framework #lemonade #unsloth #llama_cpp #hermes_agent #amd #rocm -
Ah, another gripping #ROCm blog post where we're supposed to pretend optimizing kernels on #AMD MI450 #GPUs is thrilling and not a #corporate yawn-fest 🤖. Spoiler alert: it's basically a sleep-inducing manifesto on how to decode #attention, and not the kind you need to understand this guide 🌌.
https://rocm.blogs.amd.com/software-tools-optimization/gluon-attention-decode-mi450/README.html #optimization #yawn #blogpost #decoding #HackerNews #ngated -
RT @HyperTechInvest: AMD hat ein hochmodernes, vollständig offenes Modell veröffentlicht, das ausschließlich auf MI300X- und MI325X-GPUs mit ROCm Instella-MoE trainiert wurde. Es handelt sich um ein Mixture-of-Experts-Modell mit 16 Milliarden Parametern, das pro Token nur 2,8 Milliarden Parameter aktiviert, wodurch die Inference-Kosten gesenkt werden, während es mit dichten und MoE-Modellen mit ähnlichen oder größeren aktiven Parameteranzahlen wettbewerbsfähig bleibt. Das Modell unterstützt ein 64K-Kontextfenster und wurde in sechs Phasen trainiert, von der Vorabtrainierung bis zum Reinforcement Learning. AMD führte zudem zwei neue Effizienztechniken ein: Gated Multi-head Latent Attention, das unwichtige Attention-Ausgaben filtert, sowie FarSkip-Collective, das Kommunikation und Berechnung überlappt, wodurch die Pre-Training-Geschwindigkeit um 12,7 % steigt und die Time-to-First-Token um bis zu 39,2 % reduziert wird.
mehr auf Arint.info
#AI #AMD #GPU #MachineLearning #OpenSource #ROCm #arint_info
-
Bringing PyTorch Monarch to AMD GPUs
Comments: https://news.ycombinator.com/item?id=49048689
#HackerNews #PyTorch #AMD #GPUs #Monarch #ROCm #DistributedTraining
-
#Mozilla #AI Releases #Llamafile 0.10.4, their solution for easy-to-use #LLM as a single file that work across hardware and operating systems. With Llamafile 0.10.4 is now #Transcribefile, as a new piece built off their recently announced Transcribe.cpp project
Llamafile 0.10.4 also updates against its Llama.cpp upstream build, brings a few improvements to its #Vulkan #API and #AMD #ROCm acceleration handling, HTTPS download support, and pledge/SECCOMP sandboxing support
https://www.phoronix.com/news/Llamafile-0.10.4 -
(more Linux and FOSS news in previous posts of thread)
GrapheneOS update: new exec spawning, app toggle, July 2026 security patch:
https://alternativeto.net/news/2026/7/grapheneos-update-new-exec-spawning-app-toggle-july-2026-security-patch/OPNsense 26.7 debuts on FreeBSD 15.1 with firewall, networking & future-proof enhancements:
https://alternativeto.net/news/2026/7/opnsense-26-7-debuts-on-freebsd-15-1-with-firewall-networking-and-future-proof-enhancements/FreeBSD Laptop Support Continues Improving With WiFi, GPU & Audio Driver Work:
https://www.phoronix.com/news/FreeBSD-Laptops-June-2026FreeBSD Intern Working On Porting AMD ROCm To The BSD World:
https://www.phoronix.com/news/FreeBSD-Intern-AMD-ROCmFreeBSD Desktop Installer Option Working Through NVIDIA Driver Handling, Licensing:
https://www.phoronix.com/news/FreeBSD-Desktop-Install-NVIDIAGNU Hurd Makes Progress On AArch64, Writing Translators In Rust:
https://www.phoronix.com/news/GNU-Hurd-Q2-2026scrcpy 4.1 launches with VP8/VP9 video encoding, FFmpeg 8.1, and several enhancements:
https://alternativeto.net/news/2026/7/scrcpy-4-1-launches-with-vp8-vp9-video-encoding-ffmpeg-8-1-and-several-enhancements/Mesa 26.2.0-rc1 and 26.1.5 Released: OpenCL 3.1 Support, Vulkan Updates, and Stability Fixes:
https://www.linuxcompatible.org/story/mesa-2620rc1-and-2615-released-opencl-31-support-vulkan-updates-and-stability-fixes/FastFlowLM Joins AMD to Advance AI Inference:
https://www.amd.com/en/blogs/2026/fastflowlm-joins-amd-to-advance-ai-inference.htmlROCm 7.14: TheRock Goes Production and Expands AMD’s AI Software Platform:
https://rocm.blogs.amd.com/ecosystems-and-partners/rocm-7.14-blog/README.htmlOpen Book Touch is Crowdfunding: A Buttonless, Open Hardware Answer to Kindle:
https://feed.itsfoss.com/link/24361/17379351/open-book-touch-crowdfunding#WeeklyNews #OpenSource #FOSSNews #FOSS #OpenSourceNews #News #GrapheneOS #OPNSense #FreeBSD #GNUHurd #BSD #Scrcpy #Mesa #FastFlowLM #ROCm #OpenBookTouch #FosseryTech
-
Два AMD Strix Halo в AI‑инфраструктуре: 34 контейнера на одном, ~70 tok/s Qwen3.6 на другом
На узле моей AI‑платформы крутятся 34 контейнера: Dify, RAGFlow, векторные базы, мониторинг и SSO. Большой языковой модели среди них нет: основную генерацию стек получает по LAN с DGX Spark. На втором таком же мини‑ПК я отдельно поднял локальную Qwen3.6–35B‑A3B и прогнал серию замеров от 1K до 64K при контекстном окне 256K. Обе машины — Beelink GTR9 Pro на Ryzen AI Max+ 395 (Strix Halo). Ниже — что эти коробки реально умеют: 117,4 ГиБ GTT после настройки ttm.pages_limit , p50/p95 локальных эмбеддингов и реранка, около 70 tok/s генерации через Vulkan/RADV и три грабли gfx1151.
-
#LLM performance on #AMD Radeon AI PRO R9700 (32GB VRAM) using #ROCm 7.2: input/output tokens per second and VRAM usage at a 128k context length with Q8 KV cache:
Qwen 3.6 27B ~ 168(in)/25(out) t/s, 21GB
Qwen 3.6 27B MTP ~ 127(in)/28(out) t/s, 17GB
Qwen 3.6 35B A3B ~ 246(in)/71(out) t/s, 26GB
Qwen 3.6 35B A3B MTP ~ 195(in)/74(out) t/s, 22GB
Gemma 4 31B ~ 251(in)/22 (out) t/s, 20GB
Gemma 4 26B A4B ~ 418(in)/72(out) t/s, 18GB
Gemma 4 12 B ~ 493(in)/47(out) t/s, 9GB#Ollama, Q4_K_M quantization.
-
been looking into the possibility of using AMD ROCm on the iGPU in my 5750GE box, and while it's not officially supported, I found evidence (including comments from #AMD engineers) that they are working to support older GPUs on a best-effort basis, and found this document:
https://github.com/ROCm/TheRock/blob/main/SUPPORTED_GPUS.md #ROCm
-
RT @Italianclownz: Ich habe zahlreiche Vulkan- und ROCm-Stabilitätsupdates für ROCmFP4 zusammengeführt und gepusht. Dabei habe ich einige in anderen Repositories ausstehende Stabilitätskorrekturen identifiziert, die sich auch für ROCmFP4 eigneten, und dafür gesorgt, dass die entsprechenden Autoren für ihre Codebeiträge anerkannt wurden.
mehr auf Arint.info
-
Flatpak 1.18 añade soporte para el ROCm de AMD
-
I benchmarked LLM performance on AMD Radeon AI PRO R9700 with Ollama, comparing ROCm 7.1 vs 6.4 across 8 models (Mistral, Llama, Qwen, GPT-OSS, DeepSeek, and more). Result: visible gains in prompt throughput (+87% avg) and faster responses (+11% avg). Full tables, setup, and notes in the post.
-
AMD ROCm 7.2.4 Released With Performance & Stability Fixes
AMD ROCm 7.2.4 Released With Performance & Stability Fixes #rocm #amd #videocard #ai -
AMD ROCm 7.2.4 Released With Performance & Stability Fixes
AMD ROCm 7.2.4 Released With Performance & Stability Fixes #rocm #amd #videocard #ai