#lmstudio — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #lmstudio, aggregated by home.social.
-
Today’s local LLM adventure: after updating LM Studio, my RTX 3090 dropped from 60-70 tk/s to 12-18.
Logs showed only 54 layers offloaded to GPU, despite max GPU offload and GPU-only settings. The rest was on sys CPU + RAM.
The annoying fix was switching from CUDA12 to CUDA or Vulkan. Both fixed it, with CUDA slightly faster. CUDA12 is forcing CPU offload for some reason.
TIL: if local LLM speed tanks after an update, check actual GPU layer offload and runtime.
-
Here is another thing I use containers for:
I use #lmstudio to run local llms. When you do that, you quickly notice that chatgpt and claude are not only so good because their llms are good, it's also about the tool use and the harness, even in the chat application.
A pure LLM is pretty stupid (as Mastodon people will point out "occasionally").
The real magic is via tool use and lmstudio can do that. So you need an mcp for web search, and this and that and I run all those in Containers.
-
[Перевод] Qwen3.8-27B: лучший локальный LLM, который вы, вероятно, не сможете запустить
В этом месяце Alibaba выпустила две модели, и та, о которой все писали, оказалась не той, что мы ждали. Qwen3.8-Max — это API с 2,4 триллионами параметров, и пару недель назад я с большим энтузиазмом написал о ней обзор. Но релиз, которого я действительно ждал, вышел 14 августа: Qwen3.8-27B, Apache 2.0, веса на Hugging Face, модель достаточно компактна, чтобы работать на ноутбуке. Я должен кое в чем признаться. Я не провел полный тест этой модели. Я попробовал, на своем MacBook Air M4 получал около восьми токенов в секунду и потерял терпение где-то на втором запросе. По большей части этот пост посвящен именно моей неудаче, потому что я подозреваю, что у многих из вас вечер сложится так же, как у меня. Что представляет собой Qwen3.8-27B на самом деле 27,78 миллиарда параметров, плотная модель, включающая в себя визуальный энкодер, о котором никто не объявлял заранее. Принимает на вход текст, изображения и видео. Собственный контекст — 262 144 токена, с помощью YaRN можно увеличить его примерно до миллиона, если запускать модель на сервере. Архитектура — это по-настоящему интересная часть, и это не обычный трансформер. На протяжении 64 слоёв Qwen чередует 48 слоёв Gated DeltaNet (линейное внимание) с 16 полными слоями Gated Attention в соотношении 3:1. Только эти 16 слоёв с полным вниманием имеют KV-кэш. Таким образом, на каждый токен приходится около 64 КБ, что составляет примерно четверть от объёма памяти, необходимого для обычной 64-слойной модели с плотным кодированием. Если вы запомните только одно число из этого поста, пусть это будет именно оно. Всё, что касается того, поместится ли эта модель на вашем компьютере, зависит от объёма KV-кэша.
-
🖥️ Desktop app built with #Tauri + #React, gateway binary in #Rust. Auto-detects 20 AI clients including #ClaudeCode, #Cursor, #VSCode, #Codex, #Windsurf, #Zed and #LMStudio, and writes their config for you. HTTP/OpenAPI mode connects #OpenWebUI and #n8n directly.
📦 MIT licensed, Windows & macOS signed installers, Linux in beta.
https://github.com/tsouth89/toolport -
Now that Qwen 3.8 27B released, you may want to know how to create (good) skills for your model:
https://medium.com/p/c8c735847494
#AI #ArtificialIntelligence #Laravel #PHP #Programming #Coding #Code #SoftwareDevelopment #WebDevelopment #WebDev #VibeCoding #VibeCode #Qwen #Qwen38 #LLM #HuggingFace #GGUF #Ollama #LlamaCCP #LMStudio
-
Got sucked into the AI rabbit hole a bit tonight and got LM Studio set up on my Mac mini. I have a few models downloaded plus the web search plugin activated. With a little LM Link magic, I can easily access them on my phone via the Locally app! Very impressed, so far. 🤖
-
Configuring local inference in Xcode: how I set up Xcode, OpenCode, LM Studio and Gemma 4 for local, agentic software development through the Coding Assistant UI.
https://www.patreon.com/chironcodex/posts/configuring-in-166369114
-
Hab mir im Februar ein #MacStudio mit #M4max und 128GB RAM gekauft... Für verrückte fast 4100€... Und ich dachte mir das ist der größte finanzielle Fehler ever
Hab damit dann viel #ollama #lmstudio #omlx gehostet - #LLM Experimentarium
Habs sogar ein paar Monaten nur mit den eigenen LLMs gecodet, aber Qualität naja...
Da sich die letzten Monate das #openWeights Business geändert hat und keine 120b Modelle veröffentlicht werden, die der perfekte fit dafür wären, nutze ich die 30b Modelle...
Leider mit zunehmendem Frust... #Gemma und #Qwen sind dafür einfach zu klein...
Dann bringt qwen auch nach 3.6 keine kleinen Modelle mehr raus - es gab kein 3.7 und kein 3.8 - letztes haben sie wieder man angekündigt
Bedeutet der Mac Studio war eigentlich nur als Desktop PC im Einsatz - eher #gaming und #Roblox programmierworkstation
Dann hab ich mich mal schlau gemacht - #apple hat meine 128gb Version zum frontier Modell gemacht - und die Modelle sind nirgendwo zu bekommen... Ganz merkwürdig Markt leer gesaugt
Vllt bringt Apple bald neue Mac Studio Modell raus???Okay - eBay preise abgecheckt - da gehen Modell für über 5k€ über den Tisch, aber sehr selten die 128gb Variante
eBay inseriert - 6999,99€ sofort kauf oder Preisvorschlag ab 6000€
Hab das Ding jetzt für 6200€ weiter verkauft
Leider über eBay bezahlt, deshalb wird mein Gewinn versteuert, aber ich gönne
Ist das nicht krass? Was da abgeht, das ist so verrückt auf dem Gebrauch Mac Markt!
-
Entscheidung: ich gehe auf meinem Laptop komplett von #ollama weg und nutze nur noch #lmstudio - auf dem Mac mini weiß ich noch nicht genau. LM Studio Link und die Verfügbarkeit auf iOS (App: Locally) wäre ein Grund ganz zu wechseln, aber dann gehe ich in ein geschlossenes Ökosystem.
Steht jemand vor ähnlichen Entscheidungen?
-
I'm currently testing a few #LLM #MLX runtimes/apps on my #M5 Pro 24 GB. So far, I can't say which one gets the best performance out of a local LLM like #Gemma 4 12B.
I've tested #Ollama, #LMStudio, #oMlx, and #osaurus. Each one has its own 1-2 tweaks that are supposed to squeeze out more performance.What's your experience been like?
-
While it took some finagling, I was technically able to get Deepseek v4 Flash running on my framework desktop.
Fun as a POC, but as it's slower and has a limited context window, it won't replace Qwen3.6 as my daily driver for local hosted models just yet.
-
Liebes Internet,
wenn ich die ersten Schritte in Richtung #LLMwiki mit #ClaudeCode oder alternativ lokal mit #Bionic von #LMStudio und #Gemma4 rumspielen will – und aber noch keinen Plan habe, wie ich dabei vorgehe: Welche Ressourcen (Text, Tutorials, Videos, Prompts) empfehlt ihr mir, um den Einstieg zu finden?(auf einem MacBook Air M4 und 24GB RAM)
Ziel ist, eigene Markdown Files anzusprechen und darauf basierend neue zu erstellen. Kein krasses Coding oder so
-
Everyone's paying per token to run AI. This desktop agent doesn't.
LM Studio's new Bionic runs open models — GLM 5.2 and Kimi K2.7 Code — right on your own machine. No API keys. No usage limits. Zero data retention.
Cloud when you need it. But local is the default.
The API bill was never the cost of doing AI. It was a choice.
-
best AI news yesterday, Coding Agent for your local computer: lmstudio.ai/blog/introducing-l… #LMStudio #AI #VibeCoding -
LM Studio presenta Bionic, el seu agent de IA per a models oberts.
-
WebBrain is now available on Edge too:
https://microsoftedge.microsoft.com/addons/detail/webbrain/dfbioajafcijomhljabppcelecgdgfeo
#webbrain #edge #browser #ollama #opensource #llamacpp #lmstudio #qwen #ai #gemma #automation
-
Bin gerade dabei ein wenig mit lokaler KI zu experimentieren und ich bin überrascht, wie leistungsfähig sie schon auf meinem normalen Notebook und Smartphone läuft:
https://lowmark.de/blog/lokale-ki-llm-statt-chatgpt.html
#LLM #KI #ChatGPT #AI #ollama #lmstudio #openwebui #pocketpal #Chatbot #claude -
TagSpaces 6.13 is out 🎉
The highlight for the #FOSS crowd: bring your own AI. Point TagSpaces at a local runtime — #LMStudio, #llamacpp — or any OpenAI-compatible endpoint. Run it entirely offline; your files and prompts never leave your machine.
Also in this release:
📱 New iOS app with iCloud (beta)
🤖 Rebuilt Android app on Capacitor
💾 Location backup now free
🔍 Filters for tags, locations & quick access