home.social

#rocm — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #rocm, aggregated by home.social.

fetched live
  1. I got Deepseek V4 work in Framework Desktop, an AMD Strix Halo PC with 128G memory, inside Lemonade server. I briefly tested code snippet generation, and result was pretty good. Downsides are that it's rather slow and context is short. But coming from a mini-PC it's still *very* impressive. The little Framework PC keeps giving.

    Speed was about 15 tps, but the speed stays constantly there - even when context gets longer. I remember running Kimi K2 a year ago at 3 tps!

    I tested code generation by asking it to implement an Angular module for OAUTH login client. I refined it through few iterations to add e.g. hardening and configuration features. Code quality was very good. Finally i asked it to write it out as plan.md, restarted and asked to generate code from the plan. Regenerated code was nearly identical to original round.

    The server had some instability after chat grew to about 35k long (total 15k tokens). Nothing crashes but client showed an error that stream ended prematurely. The server log showed it finished though. Overall this was good experience, with some concern about actual max content length.

    Model was unsloth/DeepSeek-V4-Flash-0731-GGUF with UD-IQ2_M quant. The Lemonade server couldn't run it out-of-the-box, complaining about unknown "Deepseek" architecture.
    - I upgraded llama.cpp to a nightly build:
    lemonade config set llamacpp.rocm_bin=b10230
    - Reloaded the llama.cpp backend from Lemonade UI.
    - File/Add model.
    - Set run parameters: --flash-attn on --reasoning on -np 1 --ctx-checkpoints 0
    - Max context: 65k

    RAM usage was at 90GB, so there is still room for another model in parallel, or better quant. No crashes, even after several hours. Next step is try Hermes Studio with the model.
    #homelab #AI #deepseek #framework #lemonade #unsloth #llama_cpp #hermes_agent #amd #rocm

  2. Ah, another gripping #ROCm blog post where we're supposed to pretend optimizing kernels on #AMD MI450 #GPUs is thrilling and not a #corporate yawn-fest 🤖. Spoiler alert: it's basically a sleep-inducing manifesto on how to decode #attention, and not the kind you need to understand this guide 🌌.
    rocm.blogs.amd.com/software-to #optimization #yawn #blogpost #decoding #HackerNews #ngated

  3. RT @HyperTechInvest: AMD hat ein hochmodernes, vollständig offenes Modell veröffentlicht, das ausschließlich auf MI300X- und MI325X-GPUs mit ROCm Instella-MoE trainiert wurde. Es handelt sich um ein Mixture-of-Experts-Modell mit 16 Milliarden Parametern, das pro Token nur 2,8 Milliarden Parameter aktiviert, wodurch die Inference-Kosten gesenkt werden, während es mit dichten und MoE-Modellen mit ähnlichen oder größeren aktiven Parameteranzahlen wettbewerbsfähig bleibt. Das Modell unterstützt ein 64K-Kontextfenster und wurde in sechs Phasen trainiert, von der Vorabtrainierung bis zum Reinforcement Learning. AMD führte zudem zwei neue Effizienztechniken ein: Gated Multi-head Latent Attention, das unwichtige Attention-Ausgaben filtert, sowie FarSkip-Collective, das Kommunikation und Berechnung überlappt, wodurch die Pre-Training-Geschwindigkeit um 12,7 % steigt und die Time-to-First-Token um bis zu 39,2 % reduziert wird.

    mehr auf Arint.info

    #AI #AMD #GPU #MachineLearning #OpenSource #ROCm #arint_info

    https://x.com/HyperTechInvest/status/2081657974364475774#m

  4. AMD Lets AI Models Optimize AMD Hardware, Undercutting CUDA's Developer-Ecosystem Advantage

    If this matters to you, share it.

    1ban.news/amd-rocm-ai-vibe-cod

    #1ban #amd #rocm #vibe #coding #tech

  5. #Mozilla #AI Releases #Llamafile 0.10.4, their solution for easy-to-use #LLM as a single file that work across hardware and operating systems. With Llamafile 0.10.4 is now #Transcribefile, as a new piece built off their recently announced Transcribe.cpp project
    Llamafile 0.10.4 also updates against its Llama.cpp upstream build, brings a few improvements to its #Vulkan #API and #AMD #ROCm acceleration handling, HTTPS download support, and pledge/SECCOMP sandboxing support
    phoronix.com/news/Llamafile-0.

  6. (more Linux and FOSS news in previous posts of thread)

    GrapheneOS update: new exec spawning, app toggle, July 2026 security patch:
    alternativeto.net/news/2026/7/

    OPNsense 26.7 debuts on FreeBSD 15.1 with firewall, networking & future-proof enhancements:
    alternativeto.net/news/2026/7/

    FreeBSD Laptop Support Continues Improving With WiFi, GPU & Audio Driver Work:
    phoronix.com/news/FreeBSD-Lapt

    FreeBSD Intern Working On Porting AMD ROCm To The BSD World:
    phoronix.com/news/FreeBSD-Inte

    FreeBSD Desktop Installer Option Working Through NVIDIA Driver Handling, Licensing:
    phoronix.com/news/FreeBSD-Desk

    GNU Hurd Makes Progress On AArch64, Writing Translators In Rust:
    phoronix.com/news/GNU-Hurd-Q2-

    scrcpy 4.1 launches with VP8/VP9 video encoding, FFmpeg 8.1, and several enhancements:
    alternativeto.net/news/2026/7/

    Mesa 26.2.0-rc1 and 26.1.5 Released: OpenCL 3.1 Support, Vulkan Updates, and Stability Fixes:
    linuxcompatible.org/story/mesa

    FastFlowLM Joins AMD to Advance AI Inference:
    amd.com/en/blogs/2026/fastflow

    ROCm 7.14: TheRock Goes Production and Expands AMD’s AI Software Platform:
    rocm.blogs.amd.com/ecosystems-

    Open Book Touch is Crowdfunding: A Buttonless, Open Hardware Answer to Kindle:
    feed.itsfoss.com/link/24361/17

    #WeeklyNews #OpenSource #FOSSNews #FOSS #OpenSourceNews #News #GrapheneOS #OPNSense #FreeBSD #GNUHurd #BSD #Scrcpy #Mesa #FastFlowLM #ROCm #OpenBookTouch #FosseryTech

  7. Два AMD Strix Halo в AI‑инфраструктуре: 34 контейнера на одном, ~70 tok/s Qwen3.6 на другом

    На узле моей AI‑платформы крутятся 34 контейнера: Dify, RAGFlow, векторные базы, мониторинг и SSO. Большой языковой модели среди них нет: основную генерацию стек получает по LAN с DGX Spark. На втором таком же мини‑ПК я отдельно поднял локальную Qwen3.6–35B‑A3B и прогнал серию замеров от 1K до 64K при контекстном окне 256K. Обе машины — Beelink GTR9 Pro на Ryzen AI Max+ 395 (Strix Halo). Ниже — что эти коробки реально умеют: 117,4 ГиБ GTT после настройки ttm.pages_limit , p50/p95 локальных эмбеддингов и реранка, около 70 tok/s генерации через Vulkan/RADV и три грабли gfx1151.

    habr.com/ru/articles/1058502/

    #strixhalo #amd #selfhosted #llm #vulkan #rocm

  8. AMD's Ryzen AI Halo turns Strix Halo into a $3,999 local AI workstation with 128GB unified memory and ROCm support. Interesting option for devs moving AI work off the cloud.

    #tech #technology #amd #ryzenai #localai #rocm

    techshowup.com/News/Article/01

  9. #LLM performance on #AMD Radeon AI PRO R9700 (32GB VRAM) using #ROCm 7.2: input/output tokens per second and VRAM usage at a 128k context length with Q8 KV cache:

    Qwen 3.6 27B ~ 168(in)/25(out) t/s, 21GB
    Qwen 3.6 27B MTP ~ 127(in)/28(out) t/s, 17GB
    Qwen 3.6 35B A3B ~ 246(in)/71(out) t/s, 26GB
    Qwen 3.6 35B A3B MTP ~ 195(in)/74(out) t/s, 22GB
    Gemma 4 31B ~ 251(in)/22 (out) t/s, 20GB
    Gemma 4 26B A4B ~ 418(in)/72(out) t/s, 18GB
    Gemma 4 12 B ~ 493(in)/47(out) t/s, 9GB

    #Ollama, Q4_K_M quantization.

  10. been looking into the possibility of using AMD ROCm on the iGPU in my 5750GE box, and while it's not officially supported, I found evidence (including comments from #AMD engineers) that they are working to support older GPUs on a best-effort basis, and found this document:

    github.com/ROCm/TheRock/blob/m #ROCm

  11. RT @Italianclownz: Ich habe zahlreiche Vulkan- und ROCm-Stabilitätsupdates für ROCmFP4 zusammengeführt und gepusht. Dabei habe ich einige in anderen Repositories ausstehende Stabilitätskorrekturen identifiziert, die sich auch für ROCmFP4 eigneten, und dafür gesorgt, dass die entsprechenden Autoren für ihre Codebeiträge anerkannt wurden.

    mehr auf Arint.info

    #AMD #LLM #OpenSource #ROCm #Stability #Vulkan #arint_info

    https://x.com/Italianclownz/status/2065900658486579430#m

  12. I benchmarked LLM performance on AMD Radeon AI PRO R9700 with Ollama, comparing ROCm 7.1 vs 6.4 across 8 models (Mistral, Llama, Qwen, GPT-OSS, DeepSeek, and more). Result: visible gains in prompt throughput (+87% avg) and faster responses (+11% avg). Full tables, setup, and notes in the post.

    meefik.dev/2026/05/31/llms-per

    #AMD #ROCm #Ollama #LLM #AI

  13. AMD ROCm 7.2.4 Released With Performance & Stability Fixes

    AMD ROCm 7.2.4 Released With Performance & Stability Fixes #rocm #amd #videocard #ai

    kbin.melroy.org/m/technology@l

  14. AMD ROCm 7.2.4 Released With Performance & Stability Fixes

    AMD ROCm 7.2.4 Released With Performance & Stability Fixes #rocm #amd #videocard #ai

    kbin.melroy.org/m/[email protected]