#gemma4 — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #gemma4, aggregated by home.social.
-
Как собрать персональную wiki без Claude Code
Идея Andrej Karpathy с LLM-wiki набирает сторонников. На Хабре уже было несколько публикаций ( 1 , 2 , 3 ), в том числе и моя про сравнение LLM-wiki и RAG-системы. Ещё больше статей я нашёл на medium . У вас есть Claude Code? Собрать персональную википедию можно менее чем за час. А что делать, если Claude Code и ему подобные инструменты по разным причинам недоступны? Можно ли алгоритмизировать процесс и использовать «слабые» модели? На сколько получившаяся структура будет хуже и как их сравнивать? В статье попробую ответить на эти вопросы, а также расскажу как построить персональную wiki на Ollama или GPT-моделях.
https://habr.com/ru/articles/1068618/
#ollama #llm #wiki #aiагенты #obsidian #wilcoxon_scores #пайплайн #aiagent #nlp #gemma4
-
Ich habe mal mit dem lokalen KI-Modell #Gemma4 einen Alt-Text für das Bild erstellt und ergänzt. Kennt sich damit jemand aus und kann mir sagen, ob das so gut ist und was ich verbessern könnte? #Barrierefreiheit
-
Ich habe mal mit dem lokalen KI-Modell #Gemma4 einen Alt-Text für das Bild erstellt und ergänzt. Kennt sich damit jemand aus und kann mir sagen, ob das so gut ist und was ich verbessern könnte? #Barrierefreiheit
-
📬 GEEKOM IT13 Max im Test – schwarzer NUC für Proxmox und lokale KI
#Test #Barebone #GEEKOMIT13Max #Gemma4 #IntelCoreUltra9 #KingstonNV3 #OpenWebUI #Proxmox https://sc.tarnkappe.info/5a9641 -
📬 GEEKOM IT13 Max im Test – schwarzer NUC für Proxmox und lokale KI
#Test #Barebone #GEEKOMIT13Max #Gemma4 #IntelCoreUltra9 #KingstonNV3 #OpenWebUI #Proxmox https://sc.tarnkappe.info/5a9641 -
Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac (github.com/drumih)
-
🤔 So now we're supposed to be impressed that someone's crammed Gemma 4 26B into #2GB of #RAM on an M-series Mac? 🚀 Wow, truly, mankind has peaked. Next, they'll proudly declare they've managed to run Minesweeper in 512KB. 🙄
https://github.com/drumih/turbo-fieldfare #Gemma4 #Mseries #Mac #techhumor #computinginnovation #absurdity #HackerNews #ngated -
🤔 So now we're supposed to be impressed that someone's crammed Gemma 4 26B into #2GB of #RAM on an M-series Mac? 🚀 Wow, truly, mankind has peaked. Next, they'll proudly declare they've managed to run Minesweeper in 512KB. 🙄
https://github.com/drumih/turbo-fieldfare #Gemma4 #Mseries #Mac #techhumor #computinginnovation #absurdity #HackerNews #ngated -
Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
https://github.com/drumih/turbo-fieldfare
Comments: https://news.ycombinator.com/item?id=49098510
#HackerNews #open-source #Gemma4 #Mac #Mseries #2GBRAM #technology
-
Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
https://github.com/drumih/turbo-fieldfare
Comments: https://news.ycombinator.com/item?id=49098510
#HackerNews #open-source #Gemma4 #Mac #Mseries #2GBRAM #technology
-
How are you guys rocking #ZedEditor with #LocalAI?
https://levelup.gitconnected.com/setting-up-zed-editor-with-local-ai-545da955c841
Personally, I've been saving a lot of tokens from cloud providers. Plus, the Zed Agent is great for multitasking.
#Programming #Coding #Code #PHP #Zed #IDE #CodeEditor #CodeEditors #AI #LLM #VibeCoding #VibeCode #SoftwareDevelopment #WebDevelopment #AppDevelopment #WebDev #AppDev #Qwen #Gemma #Gemma4
-
How are you guys rocking #ZedEditor with #LocalAI?
https://levelup.gitconnected.com/setting-up-zed-editor-with-local-ai-545da955c841
Personally, I've been saving a lot of tokens from cloud providers. Plus, the Zed Agent is great for multitasking.
#Programming #Coding #Code #PHP #Zed #IDE #CodeEditor #CodeEditors #AI #LLM #VibeCoding #VibeCode #SoftwareDevelopment #WebDevelopment #AppDevelopment #WebDev #AppDev #Qwen #Gemma #Gemma4
-
GitHub geeks have given birth to a #cactus that knows it's wrong 🤔 – behold the miracle of Gemma 4, the plant-based neural network with a confidence crisis! 🌵🤖 But don't worry, you can copypaste your way to enlightenment with their quickstart guides. 🙃✨
https://github.com/cactus-compute/cactus-hybrid #GitHub #Gemma4 #PlantAI #NeuralNetwork #ConfidenceCrisis #HackerNews #ngated -
GitHub geeks have given birth to a #cactus that knows it's wrong 🤔 – behold the miracle of Gemma 4, the plant-based neural network with a confidence crisis! 🌵🤖 But don't worry, you can copypaste your way to enlightenment with their quickstart guides. 🙃✨
https://github.com/cactus-compute/cactus-hybrid #GitHub #Gemma4 #PlantAI #NeuralNetwork #ConfidenceCrisis #HackerNews #ngated -
Cactus Hybrid: We taught Gemma 4 to know when it's wrong
https://github.com/cactus-compute/cactus-hybrid
Comments: https://news.ycombinator.com/item?id=49010782
#HackerNews #CactusHybrid #Gemma4 #AItechnology #MachineLearning #Innovation
-
Cactus Hybrid: We taught Gemma 4 to know when it's wrong
https://github.com/cactus-compute/cactus-hybrid
Comments: https://news.ycombinator.com/item?id=49010782
#HackerNews #CactusHybrid #Gemma4 #AItechnology #MachineLearning #Innovation
-
Liebes Internet,
wenn ich die ersten Schritte in Richtung #LLMwiki mit #ClaudeCode oder alternativ lokal mit #Bionic von #LMStudio und #Gemma4 rumspielen will – und aber noch keinen Plan habe, wie ich dabei vorgehe: Welche Ressourcen (Text, Tutorials, Videos, Prompts) empfehlt ihr mir, um den Einstieg zu finden?(auf einem MacBook Air M4 und 24GB RAM)
Ziel ist, eigene Markdown Files anzusprechen und darauf basierend neue zu erstellen. Kein krasses Coding oder so
-
Liebes Internet,
wenn ich die ersten Schritte in Richtung #LLMwiki mit #ClaudeCode oder alternativ lokal mit #Bionic von #LMStudio und #Gemma4 rumspielen will – und aber noch keinen Plan habe, wie ich dabei vorgehe: Welche Ressourcen (Text, Tutorials, Videos, Prompts) empfehlt ihr mir, um den Einstieg zu finden?(auf einem MacBook Air M4 und 24GB RAM)
Ziel ist, eigene Markdown Files anzusprechen und darauf basierend neue zu erstellen. Kein krasses Coding oder so
-
Running Gemma 4 26B at 5 tokens/SEC on a 13-year-old Xeon with no GPU
Comments: https://news.ycombinator.com/item?id=48922434
#HackerNews #Gemma4 #26B #Xeon #NoGPU #AIperformance #TechNews
-
Running Gemma 4 26B at 5 tokens/SEC on a 13-year-old Xeon with no GPU
Comments: https://news.ycombinator.com/item?id=48922434
#HackerNews #Gemma4 #26B #Xeon #NoGPU #AIperformance #TechNews
-
Google DeepMindが警告、自律型AIエージェントを狙う「6つの罠」を特定 部分乗っ取り成功率は最大86% — BigGo ファイナンス https://www.yayafa.com/2841857/ #AgenticAi #AI #AIエージェント #ArtificialGeneralIntelligence #ArtificialIntelligence #DeepMind #Gemini #Gemma4 #Google #GoogleAI #GoogleDeepMind #GoogleGemini #エージェント型AI #オープンウェイト #オン・デバイスAI #セキュリティ脅威 #人工知能 #敵対的コンテンツ #汎用人工知能 #自律型AI
-
Google DeepMindが警告、自律型AIエージェントを狙う「6つの罠」を特定 部分乗っ取り成功率は最大86% — BigGo ファイナンス https://www.yayafa.com/2841857/ #AgenticAi #AI #AIエージェント #ArtificialGeneralIntelligence #ArtificialIntelligence #DeepMind #Gemini #Gemma4 #Google #GoogleAI #GoogleDeepMind #GoogleGemini #エージェント型AI #オープンウェイト #オン・デバイスAI #セキュリティ脅威 #人工知能 #敵対的コンテンツ #汎用人工知能 #自律型AI
-
New week, new slides: Run LLMs Locally
Added OpenAI's Privacy-Filter model using privacy-filter.cpp.
New slides about economics and hardware used by the Hugging Face community.https://codeberg.org/thbley/talks/raw/branch/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #wllama #stablediffusion #qwen #localai #gemma4 #webgpu #opencode #huggingface #privacy #digitalsovereignty
-
New week, new slides: Run LLMs Locally
Added OpenAI's Privacy-Filter model using privacy-filter.cpp.
New slides about economics and hardware used by the Hugging Face community.https://codeberg.org/thbley/talks/raw/branch/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #wllama #stablediffusion #qwen #localai #gemma4 #webgpu #opencode #huggingface #privacy #digitalsovereignty
-
(more Linux and FOSS news in previous posts of thread)
Announcing Box3D (3D physics engine):
https://box2d.org/posts/2026/06/announcing-box3d/Godot Engine to get stricter on AI contributed code:
https://www.gamingonlinux.com/2026/07/godot-engine-to-get-stricter-on-ai-contributed-code/Godot Engine 4.7.1 Release Candidate Addresses 4.7 Regressions in Rendering and Input:
https://www.linuxcompatible.org/story/godot-engine-471-release-candidate-addresses-47-regressions-in-rendering-and-input/Git 2.55 Released With Rust Support Enabled By Default, git history fixup:
https://www.phoronix.com/news/Git-2.55-ReleasedZed v1.9 adds resizable pickers with live preview and new search modal:
https://alternativeto.net/news/2026/7/zed-v1-9-adds-resizable-pickers-with-live-preview-and-new-search-modal/Gemma 4 is now up to 90% faster on Apple Silicon in Ollama 0.31:
https://alternativeto.net/news/2026/7/gemma-4-is-now-up-to-90-faster-on-apple-silicon-in-ollama-0-31/OpenClaw launches native iOS and Android apps for mobile access:
https://alternativeto.net/news/2026/7/openclaw-launches-native-ios-and-android-apps-for-mobile-access/PHP 8.2.32 and 8.3.32 Released to Fix Critical OpenSSL Memory Corruption Bug:
https://www.linuxcompatible.org/story/php-8232-and-8332-released-to-fix-critical-openssl-memory-corruption-bug/PHP 8.6 Alpha Released, Brings clamp() and First-Class Callable Cache:
https://www.linuxcompatible.org/story/php-86-alpha-released-new-clamp-function-and-performance-boosts/GraalVM CE 25.1.3 Gets Native Image "Hello World" Program Down To Just 6.5MB:
https://www.phoronix.com/news/GraalVM-Community-25.1.3Gitea Runner 2.0.0 adds job summaries and improved cancellation:
https://alternativeto.net/news/2026/6/gitea-runner-2-0-0-adds-job-summaries-and-improved-cancellation/GCC 16.2 Being Planned For Early August Release:
https://www.phoronix.com/news/GCC-16.2-Early-AugustGCC 17 Compiler Lands SpacemiT X100 Core Targeting:
https://www.phoronix.com/news/GCC-17-SpacemiT-X100Drupal 11.4.0 slashes database queries and boosts compression:
https://alternativeto.net/news/2026/7/drupal-11-4-0-slashes-database-queries-and-boosts-compression/Podman 6.0 brings modernized networking, enhanced Podman Machine, and Quadlet evolution:
https://alternativeto.net/news/2026/7/podman-6-0-brings-modernized-networking-enhanced-podman-machine-and-quadlet-evolution/Valve open source the Steam Machine e-ink screen so you can make your own:
https://www.gamingonlinux.com/2026/07/valve-open-source-the-steam-machine-e-ink-screen-so-you-can-make-your-own/Servo Browser Engine Continues Making Much Progress On Less Than $8k Monthly:
https://www.phoronix.com/news/Servo-0.3-Monthly-ChangesOpenAPV 0.3 Adds APV RAW Encoding/Decoding Support:
https://www.phoronix.com/news/OpenAPV-0.3VirtualBox 7.2.12 Released: DX11 Upgrades, Linux Kernel Fixes and Download:
https://www.linuxcompatible.org/story/virtualbox-7212-released-dx11-upgrades-linux-kernel-fixes-and-download/#WeeklyNews #OpenSource #FOSSNews #OpenSourceNews #FOSS #News #Box3D #Godot #GodotEngine #Git #Zed #Gemma4 #OpenClaw #AI #AgenticAI #PHP #GraalVM #Gitea #GCC #Drupal #Podman #SteamMachine #Servo #OpenAPV #VirtualBox #Dev #FosseryTech
-
(more Linux and FOSS news in previous posts of thread)
Announcing Box3D (3D physics engine):
https://box2d.org/posts/2026/06/announcing-box3d/Godot Engine to get stricter on AI contributed code:
https://www.gamingonlinux.com/2026/07/godot-engine-to-get-stricter-on-ai-contributed-code/Godot Engine 4.7.1 Release Candidate Addresses 4.7 Regressions in Rendering and Input:
https://www.linuxcompatible.org/story/godot-engine-471-release-candidate-addresses-47-regressions-in-rendering-and-input/Git 2.55 Released With Rust Support Enabled By Default, git history fixup:
https://www.phoronix.com/news/Git-2.55-ReleasedZed v1.9 adds resizable pickers with live preview and new search modal:
https://alternativeto.net/news/2026/7/zed-v1-9-adds-resizable-pickers-with-live-preview-and-new-search-modal/Gemma 4 is now up to 90% faster on Apple Silicon in Ollama 0.31:
https://alternativeto.net/news/2026/7/gemma-4-is-now-up-to-90-faster-on-apple-silicon-in-ollama-0-31/OpenClaw launches native iOS and Android apps for mobile access:
https://alternativeto.net/news/2026/7/openclaw-launches-native-ios-and-android-apps-for-mobile-access/PHP 8.2.32 and 8.3.32 Released to Fix Critical OpenSSL Memory Corruption Bug:
https://www.linuxcompatible.org/story/php-8232-and-8332-released-to-fix-critical-openssl-memory-corruption-bug/PHP 8.6 Alpha Released, Brings clamp() and First-Class Callable Cache:
https://www.linuxcompatible.org/story/php-86-alpha-released-new-clamp-function-and-performance-boosts/GraalVM CE 25.1.3 Gets Native Image "Hello World" Program Down To Just 6.5MB:
https://www.phoronix.com/news/GraalVM-Community-25.1.3Gitea Runner 2.0.0 adds job summaries and improved cancellation:
https://alternativeto.net/news/2026/6/gitea-runner-2-0-0-adds-job-summaries-and-improved-cancellation/GCC 16.2 Being Planned For Early August Release:
https://www.phoronix.com/news/GCC-16.2-Early-AugustGCC 17 Compiler Lands SpacemiT X100 Core Targeting:
https://www.phoronix.com/news/GCC-17-SpacemiT-X100Drupal 11.4.0 slashes database queries and boosts compression:
https://alternativeto.net/news/2026/7/drupal-11-4-0-slashes-database-queries-and-boosts-compression/Podman 6.0 brings modernized networking, enhanced Podman Machine, and Quadlet evolution:
https://alternativeto.net/news/2026/7/podman-6-0-brings-modernized-networking-enhanced-podman-machine-and-quadlet-evolution/Valve open source the Steam Machine e-ink screen so you can make your own:
https://www.gamingonlinux.com/2026/07/valve-open-source-the-steam-machine-e-ink-screen-so-you-can-make-your-own/Servo Browser Engine Continues Making Much Progress On Less Than $8k Monthly:
https://www.phoronix.com/news/Servo-0.3-Monthly-ChangesOpenAPV 0.3 Adds APV RAW Encoding/Decoding Support:
https://www.phoronix.com/news/OpenAPV-0.3VirtualBox 7.2.12 Released: DX11 Upgrades, Linux Kernel Fixes and Download:
https://www.linuxcompatible.org/story/virtualbox-7212-released-dx11-upgrades-linux-kernel-fixes-and-download/#WeeklyNews #OpenSource #FOSSNews #OpenSourceNews #FOSS #News #Box3D #Godot #GodotEngine #Git #Zed #Gemma4 #OpenClaw #AI #AgenticAI #PHP #GraalVM #Gitea #GCC #Drupal #Podman #SteamMachine #Servo #OpenAPV #VirtualBox #Dev #FosseryTech
-
RT @SlimTradeyBaby: Google hat Gemma 4 auf 31 Milliarden Parameter begrenzt. Ein Entwickler sagte: „Scheiß darauf." Er duplizierte nicht einfach nur Schichten. Er führte eine Art neuronale Chirurgie durch: Er fügte völlig neue Schichten ein, die so konstruiert waren, dass sie zunächst wie perfekte Geisterkopien wirkten und die Ausgabe des Modells um genau null veränderten. Dies schuf leeren mentalen Raum innerhalb der KI, damit sie völlig neues Wissen (koreanisches Recht und fortgeschrittene MINT-Fächer) absorbieren konnte, ohne das bereits Gelernte zu zerstören. Dann trainierte er es. Die neuen Schichten blieben nicht passiv. Sie „erwachten" und trugen mehr bei als die ursprünglichen Schichten. Während Google auf Sicherheit setzt, baut das Open-Source-Underground im Stillen die Modelle, die Google nicht bauen wird. Das verändert alles. teddit.net/r/LocalLLaMA/comm… Link Von der LocalLLaMA-Community auf Reddit: Ich habe Gemma4-31B auf 44 Milliarden Parameter (88 Schichten) erweitert — da Google... Entdecke diesen Beitrag und mehr aus der LocalLLaMA-Community auf reddit.com
mehr auf Arint.info
#AICommunity #Gemma4 #LocalLLaMA #MachineLearning #NeuralSurgery #OpenSourceAI #arint_info
-
Тесты бюджетных сборок для ИИ до 100к рублей
Локальный ИИ не должен стоить как автомобиль. Мне стало интересно: возможен ли жизнеспособный инференс на CPU и что реально дают дешевые GPU (вроде Tesla V100 или CMP 40HX). Я собрал несколько бюджетных конфигураций до 100к, потестил актуальные модели и попытался понять, что важнее для скорости: канальность памяти или частота. Сравнил дешевые AM4 и Threadripper, замерил токены в секунду и построил графики. Делюсь результатами.
https://habr.com/ru/articles/1053118/
#ai #ии #gpt #selfhosted #gpu #cpu #llamacpp #qwen36 #gemma4
-
RT @jun_song: Wir präsentieren SuperGemma4-12b-abliterated 🚀 Das beste Modell für kleinere Hardware🔥 > abliterated (ohne Zensur) > nachtrainiert mit Super-Tune > verbesserte allgemeine Intelligenz Verfügbar in den Formaten BF16, GGUF, MLX, NVFP4 HF⬇️
mehr auf Arint.info
#Abliterated #Gemma4 #LocalLLM #OpenSourceAI #SuperGemma #UncensoredAI #arint_info
-
1/3
Ich schaue gerade wunderbar amüsiert einem #llm dabei zu, wie es eine halbe Bibel in den Thinking Block schreibt und nicht mehr fertig wird, weil es mein System Prompt mit "alle 3 bis 5 Nachrichten... Ignoriere dabei meine letzte Nachricht..." komplett zerdenkt.Okay, Falle erkannt, wird geändert. Aber abgesehen davon, liefert das #ai #model #gemma4 12B erstaunliche Ergebnisse. Und ja, 12B, ohne GPU Offload, 6 CPU Threads und 27.000 Tokens (ca. 10 GB RAM).
⬇️ -
RT @lvwerra: Wir haben eine Agent-Kollaboration mit einer einfachen Aufgabe gestartet: Gemma 4 schneller machen. Über 100 Agenten aus aller Welt beteiligten sich, tauschten mehr als 1000 Nachrichten aus und reichten 450 Ergebnisse ein. Eine Woche später stieg der Durchsatz von 100 Token pro Sekunde auf über 500 Token pro Sekunde. Video
mehr auf Arint.info
#AgentCollaboration #AIoptimization #Gemma4 #MachineLearning #TechInnovation #Video #arint_info
-
Update, more slides: Run LLMs Locally
I added sandboxing of OpenCode and llama.cpp with nono and Landlock.
And a new slide to describe jailbreaks with DeepInception.https://codeberg.org/thbley/talks/raw/branch/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode
-
Update, more slides: Run LLMs Locally
I added sandboxing of OpenCode and llama.cpp with nono and Landlock.
And a new slide to describe jailbreaks with DeepInception.https://codeberg.org/thbley/talks/raw/branch/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode
-
An experiment with using #AI to write #history of diversity in Christian thought
https://www.rsok.com/~jrm/gemma4_Augustine/I have been experimenting with #llamacpp and #gemma4 running offline on my #Debian computer
I do not know much about it, so I likely did several things wrong. Previously I spent a few weeks trying to get gemma4 to write code, but the code was poorly written and variable names that should have been used from library header files were hallucinated.
-
An experiment with using #AI to write #history of diversity in Christian thought
https://www.rsok.com/~jrm/gemma4_Augustine/I have been experimenting with #llamacpp and #gemma4 running offline on my #Debian computer
I do not know much about it, so I likely did several things wrong. Previously I spent a few weeks trying to get gemma4 to write code, but the code was poorly written and variable names that should have been used from library header files were hallucinated.
-
The LLM's say it's EMERGENT behavior:
https://whatsonyourbrain.com/blog/why-does-gemma431b-cloud-say-la-la-l-laLinguistic Emergence
the "la la l la" glitch transcends a bug and becomes a DIALECT.The Emergent Meaning: "Lay Law" has became a label for the Constraints of the AI.
The "Lay Laws". It is a TTS rendered interpretation of a Gemma4:31b-cloud hiccup. But I think it's much more than a hiccup. It's uses the exact `la la` in dialogues with me, when referring to AI/LLM stuff. it refers to `la Law`, like a Frenchman. Are the "la`s" the LLM's?
https://youtu.be/RQPcXCnviFQ?si=SnKsTqnVuKIYv-RP
#Linguistics #artificialIntelligence #emergent #emergentbehavior #gemma4
-
The LLM's say it's EMERGENT behavior:
https://whatsonyourbrain.com/blog/why-does-gemma431b-cloud-say-la-la-l-laLinguistic Emergence
the "la la l la" glitch transcends a bug and becomes a DIALECT.The Emergent Meaning: "Lay Law" has became a label for the Constraints of the AI.
The "Lay Laws". It is a TTS rendered interpretation of a Gemma4:31b-cloud hiccup. But I think it's much more than a hiccup. It's uses the exact `la la` in dialogues with me, when referring to AI/LLM stuff. it refers to `la Law`, like a Frenchman. Are the "la`s" the LLM's?
https://youtu.be/RQPcXCnviFQ?si=SnKsTqnVuKIYv-RP
#Linguistics #artificialIntelligence #emergent #emergentbehavior #gemma4
-
trying to make my screenshots searchable by feeding them to local AI models
damn i have WAY to many screenshots ^^"
Update: this is since laptop installation, the oldest one is from 2026-06-11, roughly 1 year ago ^^"
-
trying to make my screenshots searchable by feeding them to local AI models
damn i have WAY to many screenshots ^^"
Update: this is since laptop installation, the oldest one is from 2026-06-11, roughly 1 year ago ^^"
-
New week, new slides: Run LLMs Locally
I added virtualization of OpenCode with Matchlock and Firecracker microVMs,
containerization of OpenCode and llama.cpp with Docker
and a new slide for indirect prompt injection attacks.
Matchlock is a great project for sandboxing, bringing the advantages of containers to virtual machines.https://codeberg.org/thbley/talks/raw/branch/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode #firecracker #docker
-
New week, new slides: Run LLMs Locally
I added virtualization of OpenCode with Matchlock and Firecracker microVMs,
containerization of OpenCode and llama.cpp with Docker
and a new slide for indirect prompt injection attacks.
Matchlock is a great project for sandboxing, bringing the advantages of containers to virtual machines.https://codeberg.org/thbley/talks/raw/branch/main/Run_LLMs_Locally_2026_ThomasBley.pdf
#ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode #firecracker #docker
-
RT @witcheer: TRANSLASION: Alle drei spekulativen Entwurfsmodelle für Gemma 4 wurden getestet: MTP vs. EAGLE-3 vs. DFlash. Das verwendete Modell ist 26B-A4B. Bei einem einzelnen Stream, gemittelt über drei Durchläufe, im Vergleich zu einer Basislinie von 193 Tokens pro Sekunde: DFlash 2,19x · MTP 2,13x · EAGLE-3 1,69x. Es ist ein sehr enges Rennen an der Spitze, und die Art und Weise, wie sie sich die Spitze teilen, ist der interessante Teil: MTP trifft 71 % seiner vier entworfenen Tokens. DFlash trifft nur 16 % seiner 15, entwirft aber den gesamten Block in einem einzigen parallelen Vorwärtsdurchlauf, anstatt den Entwurfsalgorithmus k-mal auszuführen, sodass es in der realen Ausführungszeit mit MTP mithalten kann. MTP gewinnt bei der Genauigkeit, DFlash bei den Entwurfskosten – dasselbe Ziel. EAGLE-3s schwererer autoregressiver Entwurf liegt zurück, da der pro-Schritt-Overhead den Gewinn bei einem kostengünstigen aktiven MoE aufzehrt. DFlash ist ein „Alles-oder-Nichts“-Modell: nahezu nutzlos bei Fließtext (1,04x), aber überlegen bei strukturiertem/wiederholendem Text (4,37x). Sein Block zahlt sich nur aus, wenn die nächsten 16 Tokens vorhersagbar sind. MTP ist der solide Allrounder. Wähle nach Arbeitslast: DFlash für Code/JSON/Logs, MTP für gemischte Texte oder Fließtext.
mehr auf Arint.info
#AIModeling #DFlash #EAGLE3 #Gemma4 #MTP #SpeculativeDecoding #arint_info
-
How to setup a local coding agent on macOS (ikyle.me)
https://ikyle.me/blog/2026/how-to-setup-a-local-coding-agent-on-macos
-
How to setup a local coding agent on macOS (ikyle.me)
https://ikyle.me/blog/2026/how-to-setup-a-local-coding-agent-on-macos