home.social

#vram — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #vram, aggregated by home.social.

fetched live
  1. GTX 1080 Ti не сдаётся: выбираем лучшую MoE-модель для llama-server среди 35B, 120B и 176B

    Привет, Хабр! В прошлой статье я рассказывал, как запустить 176B-класс Qwen3.8-Flash-Next на GTX 1080 Ti, и тогда это был, скорее, эксперимент на грани возможного. Но недавно я наткнулся на материал на ast-softpro.ru , где ребята запустили Qwen 3.8-27B на двух RTX 5060 Ti по 16 ГБ (32 ГБ суммарно). Если коротко, они подтвердили, что гибридная архитектура DeltaNet + attention позволяет эффективно работать с длинным контекстом, а на их стенде модель в квантизации Q4_K_M выдавала около 24 токенов/с. Именно это, а также комментарии на Хабре, подтолкнуло меня к серии экспериментов, о которых пойдёт речь. Забегу наперед, я был удивлён результату: на моём стенде, который я опишу ниже, результаты на одной из моделей, которая как минимум не хуже, оказались сопоставимыми.

    habr.com/ru/articles/1086998/

    #ai #llmмодели #llamacpp #server #эксперимент #старое_компьютерное_железо #видеокарта #gtx_1080 #vram

  2. He leído en los foros de #DiabloIV (en #Steam) que lo de la #VRAM no es un problema exclusivo de #GNU con #Linux. Vamos, que le pasa a todo el mundo. Por eso, y otras cosas, se lo están fundiendo en las notas del juego. Al parecer no lo arreglan.

  3. Что на самом деле происходит, когда видеопамять заканчивается

    В сентябре 2026-го разводить холивар о том, хватает ли 8 ГБ видеопамяти, уже как‑то неприлично. И даже не потому что вопрос закрыт, а потому что выбора особо не осталось. Ну, посудите сами. RTX 50 Super, которую нам два года обещали как спасение с 3-гигабайтными чипами GDDR7, на Gamescom похоронили окончательно. А то, что осталось на полках, дорожает на глазах: за один только август

    habr.com/ru/companies/x-com/ar

    #xcomshop #vram #pcie_30 #гейминг

  4. Как я выбирал локальную модель для кодинга на 16 ГБ VRAM — и трижды менял победителя

    Сначала я выбрал модель, которая отлично писала код. Подключил её к OpenCode и выяснил, что создавать файлы на диске она просто не умеет. Взял другую — та уверенно работала с файлами, но полностью поплыла на русском языке. ​В итоге победила модель, которую я поначалу оставил «для экспериментов»: слишком большая, на грани по квантованию и, казалось, сплошной компромисс. ​Ниже — о том, как я подбирал рабочую лошадку под WordPress и фронтенд на видеокарте с 16 ГБ памяти. Со скоростями, квантами и неожиданным экзаменом по русскому языку.

    habr.com/ru/articles/1083490/

    #локальные_LLM #локальные_нейросети #Qwen3Coder #llamacpp #OpenCode #модели_для_программирования #квантование_моделей #VRAM #WordPressразработка #агентское_программирование

  5. DeepSeek V4.1 Flash activates ~8B parameters per token—but the checkpoint is still ~510GB.

    The important improvement is elsewhere: DeepSeek reports an 890-byte/token global KV cache, about ¼ of V4 Flash, which helps long-context serving.

    But “8B active” does not mean “8B local model.” You still have to store the 552B backbone, ~196.6B Engram memory, and ~14B draft module.

    Active compute ≠ model footprint.

    #DeepSeek #LocalAI #MoE #VRAM #KVCache

  6. RT @tomshardware: Eine von China modifizierte Nvidia RTX 5090 mit massiven 96 GB Speicher erscheint auf Alibaba für weniger als 4.000 Dollar – 3x mehr VRAM zum 65%igen Preis des Originals.

    mehr auf Arint.info

    #Alibaba #ChinaTech #Grafikkarte #Nvidia #RTX5090 #VRAM #arint_info

    https://x.com/tomshardware/status/2098451092103459164

  7. Я прогнал топ VRAM-калькуляторов для LLM — все ошибаются в одном и том же

    Наивная формула «веса + KV < VRAM» промахивается по двум причинам, и обе стоят денег. Первая: движок (vLLM, SGLang, TensorRT-LLM) резервирует почти всю память под свой пул ещё до того, как вы посчитали KV — поэтому ответ «сколько запросов влезет» всегда мимо, в сторону «влезет», а потом OOM под нагрузкой. Вторая, которую в уме не посчитать: на DeepSeek-V2-Lite обычная формула завышает KV в 7–11 раз, и промах уже в обратную сторону — зря откажетесь ставить модель. Я прогнал топ VRAM-калькуляторов из выдачи Google, показал на скриншотах, где ломается каждый (даже лучший), сверил числа с живым vLLM на A100/H100 — расхождение ~1% — и собрал ridgepoint: калькулятор, который считает как движок, а не как учебник. Ставится локально ( pip install ridgepoint ), GPU не нужен. Сколько реально влезет

    habr.com/ru/articles/1079788/

    #llm #inference #vllm #vram #kvcache #gpu #machine_learning #python #rust #sglang

  8. Почему профессиональная GPU стоит впятеро дороже игровой с той же начинкой

    Привет, Хабр! На связи Илья Мартысь из Рег.облака.

    habr.com/ru/companies/runity/a

    #видеокарты #облачные_сервисы #выделенные_серверы #драйвер #железо #память #gpu #vram #облако #облачные_технологии

  9. Селф-хост Ornith-1.5 и Qwen3.8: тестирую 18 квантов на одной карте и похоже не покупаю RTX5090

    Встала задача - срочно за вечер решить практический вопрос: какой квант ставить на карту и нужна ли мне RTX5090 домой (спойлер: нет, 9B и Q5/Q6 и я буду счастлив). Прогнал 18 конфигураций по схеме - три модели, пять-шесть квантов на каждую, один и тот же набор из 22 реальных задач.

    habr.com/ru/articles/1074434/

    #ии #квантизация #llm #selfhosted #vram #rtx_pro_6000

  10. السّلام عليكم
    ليوم في
    #Linux, #Privacy, Open-source and open-web #news
    Linux 7.3
    حسب الأخبار الحاليّة ينجّم يحسّن إستعمالو على الماكينات "البطاطا"
    "Potato" #PCs
    بالتّحديد في مجال ال
    #Gaming
    بتحسين التّعامل مع ال
    #VRAM
    الشّويّة
    pixelcluster.dev/VRAM-Overcomm
    الوحيد إلّي يحبّوا يجرّبوا
    #Linux
    ماغير ما يعملوا المرج متاع صبّانو، ثمّة حاجة إسمها
    #WSL (Windows Subsystem for Linux)
    إلّي يخلّيك تجرّب برشة
    #Distributions
    على
    windows
    منهم
    #Ubuntu
    إلّي تشهد نموّ أكثر من إستعمالو بالطّريقة التّقليديّة
    windowslatest.com/2026/08/16/u

  11. Qwen 35B: 200 tokens por segundo em 8GB de VRAM? 🚀

    Imagina rodar um modelo de IA com 150 a 200 tokens por segundo na sua placa de 8GB de VRAM? 🤯

    O Qwen 3.8 35B híbrido tá chegando e o Alibaba já começou a atualizar os repositórios! Enquanto isso, o Qwen 3.6 27B denso já roda rápido pra caramba — e o 35B vem pra ser 5 a 6 vezes mais rápido. 🔥

    ⚡ O que esperar:
    - Modelo 35B com apenas 3B de parâmetros ativos
    - De 150 a 200 tokens por segundo em hardware...

    #Qwen #IA #VRAM #MorningCrypto

  12. hw-smi v1.6 brings support for data logging!

    You have requested an option to log the telemetry data (#GPU​/VRAM usage, #VRAM bandwidth, temperature, power, fan speed, #PCIe bandwidth etc.) from hw-smi to a file. Today I have implemented exactly that. Have fun monitoring your hardware and applications, be it #Intel Arc (Pro) or any other #Nvidia​/​#AMD GPUs on Windows and #Linux! 🖖

    github.com/ProjectPhysX/hw-smi

  13. 🎉 #Linux 7.3: Now celebrating the revolutionary concept of "Don't run out of #VRAM, or else!" 🙃 Apparently, the pinnacle of #innovation is making sure you don't overpromise your virtual memory—truly groundbreaking stuff for the #gaming elite. 🚀
    pixelcluster.dev/VRAM-Overcomm #7.3 #LinuxNews #HackerNews #ngated

  14. NVIDIA sold the RTX 5060 Ti with 8GB or 16GB. Same chip. Launch price difference: $50.
    Today the market difference is $191.
    The 4060 Ti ran the same experiment two years earlier and landed within 50 cents per gigabyte.

    buysellram.com/blog/nvidia-con

    #NVIDIA #GPUPrices #RTX50 #RTX5090 #VRAM #GDDR7 #DRAM #GPU #PCHardware #ITAD #ITAssetDisposition #DataCenter #ResaleValue #technology

  15. NVIDIA sold the RTX 5060 Ti with 8GB or 16GB. Same chip. Launch price difference: $50.
    Today the market difference is $191.
    The 4060 Ti ran the same experiment two years earlier and landed within 50 cents per gigabyte.

    buysellram.com/blog/nvidia-con

    #NVIDIA #GPUPrices #RTX50 #RTX5090 #VRAM #GDDR7 #DRAM #GPU #PCHardware #ITAD #ITAssetDisposition #DataCenter #ResaleValue

  16. buysellram.com/blog/nvidia-con

    Two findings from the August 2026 NVIDIA GPU price data.

    First: memory capacity rather than performance tier is setting the premium over MSRP. Nothing with 8GB or 12GB lists more than 13% above launch price. Nothing with 16GB or more lists less than 18% above it. The band between is empty — and the RTX 5060 Ti 16GB, the cheapest card in the upper group, carries a larger premium than the RTX 5080.

    The cleanest evidence comes from the pairs where the same GPU shipped in two memory configurations. On the RTX 5060 Ti, the extra 8GB costs $191 in the market: $23.88 per gigabyte, against $6.25 at launch and roughly $10 for the GDDR7 modules themselves. The RTX 4060 Ti, launched two years earlier, lands at $24.38.

    Second: across the RTX 50 line, used cards now list 0–14% below new, three of them at zero. The RTX 40 generation, measured the same day, still lists 11–36% below.

  17. NVIDIA sold the RTX 5060 Ti with 8GB or 16GB. Same chip. Launch price difference: $50.
    Today the market difference is $191.
    The 4060 Ti ran the same experiment two years earlier and landed within 50 cents per gigabyte.

    buysellram.com/blog/nvidia-con

    #NVIDIA #GPUPrices #RTX50 #RTX5090 #VRAM #GDDR7 #DRAM #GPU #PCHardware #ITAD #ITAssetDisposition #DataCenter #ResaleValue #tech

  18. NVIDIA RTX PRO 5000 Blackwell с 72 Гб видеопамяти. Есть ли смысл переплачивать за «половинку» флагмана?

    RTX PRO 5000 Blackwell на 72 Гб: золотая середина для локальных ИИ-моделей или переоцененный апгрейд? Разбираемся в нашем обзоре.

    habr.com/ru/companies/hostkey/

    #NVIDIA #Blackwell #RTX_PRO_5000 #GPU #VRAM #LLM #инференс #DeepSeek #CUDA #hostkey

  19. RT @LLMJunky: Ich werde die Kimi K3-Gewichte NICHT herunterladen. Hauptsächlich, weil ich nicht über 47 TB VRAM verfüge, um es zu betreiben, und in 3 Wochen etwas Besseres verfügbar sein wird, lol.

    mehr auf Arint.info

    #AI #GPU #KimiK3 #MachineLearning #TechHumor #VRAM #arint_info

    https://x.com/LLMJunky/status/2081857661524377920#m

  20. RT @LLMJunky: Ich werde die Kimi K3-Modelle NICHT herunterladen. Vor allem, weil ich nicht 47 TB VRAM habe, um sie zu betreiben, und in 3 Wochen etwas Besseres verfügbar sein wird, lol.

    mehr auf Arint.info

    #AI #DeepLearning #KimiK3 #MachineLearning #VRAM #arint_info

    https://x.com/LLMJunky/status/2081857661524377920#m

  21. 💻🎮 Ah, the RTX 2080 Ti memory upgrade! Because who doesn't want to spend more on an ancient card instead of just buying a new one? 🙄 Double the #VRAM, double the opportunity to brag about your retro tech expertise at your next LAN party. 😂
    gpusolutions.net/rbservices/gr #RTX2080Ti #MemoryUpgrade #RetroTech #LANParty #GamingHumor #HackerNews #ngated

  22. Tip I keep giving people about that I wonder if anyone has a counterpoint too...

    A is still a Graphics Card, even if it's old or low end. Even the obsolete 1050 in my laptop can render billions of polygons at 60fps, more if I give it external cooling, and is capable of every major shader operation, short of , that a newer card is: other than a potentially heavy short-term load like and can throw when you're building something complicated, the only useful thing my cards are short on is . Hell, even the embedded in the old i7 , and the in my desktop's 7600 can run less intensive games like Black Mesa without many issues. (I know this because I've run them with the graphics card deactivated, I actually can't get Black Mesa to run with it.)

    If you're building a computer and don't know what to get, just get the cheapest graphics card that meets your VRAM needs. (Probably at least 8GB with how unoptimized User Generated Content in spaces like Resonite can get, or how unoptimized "professional" developers have gotten)

    Outside of specific issues, there's really not much need to get the high tier or later cards, unless you know a specific need for it: the 5 year old mid-tier will give you a passable experience in the majority of cases, and will likely have fewer driver/compatibility issues anyway. Newer and higher tier is just not worth another few hundred dollars these days.

    Someone else may have a reason I'm wrong though: I'm open to a reason this is bad advice.

  23. The #OpenCode is working. Check when the Nvidia is used for it
    ⎿  [Image #1]

    ● I'll check how #Ollama uses the #Nvidia GPU.

    ● Bash(nvidia-smi --query-gpu=name,memory.used,memory.total,utilization.gpu --format=csv,noheader 2>/dev/null)
    ⎿  NVIDIA GeForce RTX 2060, 952 MiB, 6144 MiB, 9 %

    ● Bash(ollama ps 2>/dev/null)
    ⎿  NAME ID SIZE PROCESSOR UNTIL
    qwen2.5-coder:7b-16k 0533dd01f237 7.2 GB 30%/70% CPU/GPU 4 minutes from now

    ● The model is currently loaded and actively using the GPU. Here's the breakdown:

    - Model: #qwen2.5-coder:7b-16k — 7.2 GB total
    - Split: 30% CPU / 70% GPU — the model is split across RAM and VRAM because it doesn't fit entirely in the RTX 2060's 6 GB VRAM
    - Current #GPU #VRAM usage: 952 MiB used of 6144 MiB (the rest of the model layers sit in system RAM)
    - GPU compute utilization: 9% right now (idle between requests)

    The GPU is used whenever OpenCode sends a prompt to Ollama — during inference, 70% of the model layers run on the #RTX2060 and 30% on the #CPU. The model stays loaded in VRAM for 4 more minutes of idle time before Ollama unloads it.

    #LocalLLM

Share on Mastodon

Enter the server where you have an account.