home.social

#gemma4 — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #gemma4, aggregated by home.social.

fetched live
  1. Как собрать персональную wiki без Claude Code

    Идея Andrej Karpathy с LLM-wiki набирает сторонников. На Хабре уже было несколько публикаций ( 1 , 2 , 3 ), в том числе и моя про сравнение LLM-wiki и RAG-системы. Ещё больше статей я нашёл на medium . У вас есть Claude Code? Собрать персональную википедию можно менее чем за час. А что делать, если Claude Code и ему подобные инструменты по разным причинам недоступны? Можно ли алгоритмизировать процесс и использовать «слабые» модели? На сколько получившаяся структура будет хуже и как их сравнивать? В статье попробую ответить на эти вопросы, а также расскажу как построить персональную wiki на Ollama или GPT-моделях.

    habr.com/ru/articles/1068618/

    #ollama #llm #wiki #aiагенты #obsidian #wilcoxon_scores #пайплайн #aiagent #nlp #gemma4

  2. Ich habe mal mit dem lokalen KI-Modell #Gemma4 einen Alt-Text für das Bild erstellt und ergänzt. Kennt sich damit jemand aus und kann mir sagen, ob das so gut ist und was ich verbessern könnte? #Barrierefreiheit

  3. Ich habe mal mit dem lokalen KI-Modell #Gemma4 einen Alt-Text für das Bild erstellt und ergänzt. Kennt sich damit jemand aus und kann mir sagen, ob das so gut ist und was ich verbessern könnte? #Barrierefreiheit

  4. 🤔 So now we're supposed to be impressed that someone's crammed Gemma 4 26B into #2GB of #RAM on an M-series Mac? 🚀 Wow, truly, mankind has peaked. Next, they'll proudly declare they've managed to run Minesweeper in 512KB. 🙄
    github.com/drumih/turbo-fieldf #Gemma4 #Mseries #Mac #techhumor #computinginnovation #absurdity #HackerNews #ngated

  5. 🤔 So now we're supposed to be impressed that someone's crammed Gemma 4 26B into #2GB of #RAM on an M-series Mac? 🚀 Wow, truly, mankind has peaked. Next, they'll proudly declare they've managed to run Minesweeper in 512KB. 🙄
    github.com/drumih/turbo-fieldf #Gemma4 #Mseries #Mac #techhumor #computinginnovation #absurdity #HackerNews #ngated

  6. GitHub geeks have given birth to a #cactus that knows it's wrong 🤔 – behold the miracle of Gemma 4, the plant-based neural network with a confidence crisis! 🌵🤖 But don't worry, you can copypaste your way to enlightenment with their quickstart guides. 🙃✨
    github.com/cactus-compute/cact #GitHub #Gemma4 #PlantAI #NeuralNetwork #ConfidenceCrisis #HackerNews #ngated

  7. GitHub geeks have given birth to a #cactus that knows it's wrong 🤔 – behold the miracle of Gemma 4, the plant-based neural network with a confidence crisis! 🌵🤖 But don't worry, you can copypaste your way to enlightenment with their quickstart guides. 🙃✨
    github.com/cactus-compute/cact #GitHub #Gemma4 #PlantAI #NeuralNetwork #ConfidenceCrisis #HackerNews #ngated

  8. Liebes Internet,
    wenn ich die ersten Schritte in Richtung #LLMwiki mit #ClaudeCode oder alternativ lokal mit #Bionic von #LMStudio und #Gemma4 rumspielen will – und aber noch keinen Plan habe, wie ich dabei vorgehe: Welche Ressourcen (Text, Tutorials, Videos, Prompts) empfehlt ihr mir, um den Einstieg zu finden?

    (auf einem MacBook Air M4 und 24GB RAM)

    Ziel ist, eigene Markdown Files anzusprechen und darauf basierend neue zu erstellen. Kein krasses Coding oder so

    #followerpower #ai #ki

  9. Liebes Internet,
    wenn ich die ersten Schritte in Richtung #LLMwiki mit #ClaudeCode oder alternativ lokal mit #Bionic von #LMStudio und #Gemma4 rumspielen will – und aber noch keinen Plan habe, wie ich dabei vorgehe: Welche Ressourcen (Text, Tutorials, Videos, Prompts) empfehlt ihr mir, um den Einstieg zu finden?

    (auf einem MacBook Air M4 und 24GB RAM)

    Ziel ist, eigene Markdown Files anzusprechen und darauf basierend neue zu erstellen. Kein krasses Coding oder so

    #followerpower #ai #ki

  10. New week, new slides: Run LLMs Locally

    Added OpenAI's Privacy-Filter model using privacy-filter.cpp.
    New slides about economics and hardware used by the Hugging Face community.

    codeberg.org/thbley/talks/raw/

    #ai #llm #llamacpp #wllama #stablediffusion #qwen #localai #gemma4 #webgpu #opencode #huggingface #privacy #digitalsovereignty

  11. New week, new slides: Run LLMs Locally

    Added OpenAI's Privacy-Filter model using privacy-filter.cpp.
    New slides about economics and hardware used by the Hugging Face community.

    codeberg.org/thbley/talks/raw/

    #ai #llm #llamacpp #wllama #stablediffusion #qwen #localai #gemma4 #webgpu #opencode #huggingface #privacy #digitalsovereignty

  12. (more Linux and FOSS news in previous posts of thread)

    Announcing Box3D (3D physics engine):
    box2d.org/posts/2026/06/announ

    Godot Engine to get stricter on AI contributed code:
    gamingonlinux.com/2026/07/godo

    Godot Engine 4.7.1 Release Candidate Addresses 4.7 Regressions in Rendering and Input:
    linuxcompatible.org/story/godo

    Git 2.55 Released With Rust Support Enabled By Default, git history fixup:
    phoronix.com/news/Git-2.55-Rel

    Zed v1.9 adds resizable pickers with live preview and new search modal:
    alternativeto.net/news/2026/7/

    Gemma 4 is now up to 90% faster on Apple Silicon in Ollama 0.31:
    alternativeto.net/news/2026/7/

    OpenClaw launches native iOS and Android apps for mobile access:
    alternativeto.net/news/2026/7/

    PHP 8.2.32 and 8.3.32 Released to Fix Critical OpenSSL Memory Corruption Bug:
    linuxcompatible.org/story/php-

    PHP 8.6 Alpha Released, Brings clamp() and First-Class Callable Cache:
    linuxcompatible.org/story/php-

    GraalVM CE 25.1.3 Gets Native Image "Hello World" Program Down To Just 6.5MB:
    phoronix.com/news/GraalVM-Comm

    Gitea Runner 2.0.0 adds job summaries and improved cancellation:
    alternativeto.net/news/2026/6/

    GCC 16.2 Being Planned For Early August Release:
    phoronix.com/news/GCC-16.2-Ear

    GCC 17 Compiler Lands SpacemiT X100 Core Targeting:
    phoronix.com/news/GCC-17-Space

    Drupal 11.4.0 slashes database queries and boosts compression:
    alternativeto.net/news/2026/7/

    Podman 6.0 brings modernized networking, enhanced Podman Machine, and Quadlet evolution:
    alternativeto.net/news/2026/7/

    Valve open source the Steam Machine e-ink screen so you can make your own:
    gamingonlinux.com/2026/07/valv

    Servo Browser Engine Continues Making Much Progress On Less Than $8k Monthly:
    phoronix.com/news/Servo-0.3-Mo

    OpenAPV 0.3 Adds APV RAW Encoding/Decoding Support:
    phoronix.com/news/OpenAPV-0.3

    VirtualBox 7.2.12 Released: DX11 Upgrades, Linux Kernel Fixes and Download:
    linuxcompatible.org/story/virt

    #WeeklyNews #OpenSource #FOSSNews #OpenSourceNews #FOSS #News #Box3D #Godot #GodotEngine #Git #Zed #Gemma4 #OpenClaw #AI #AgenticAI #PHP #GraalVM #Gitea #GCC #Drupal #Podman #SteamMachine #Servo #OpenAPV #VirtualBox #Dev #FosseryTech

  13. (more Linux and FOSS news in previous posts of thread)

    Announcing Box3D (3D physics engine):
    box2d.org/posts/2026/06/announ

    Godot Engine to get stricter on AI contributed code:
    gamingonlinux.com/2026/07/godo

    Godot Engine 4.7.1 Release Candidate Addresses 4.7 Regressions in Rendering and Input:
    linuxcompatible.org/story/godo

    Git 2.55 Released With Rust Support Enabled By Default, git history fixup:
    phoronix.com/news/Git-2.55-Rel

    Zed v1.9 adds resizable pickers with live preview and new search modal:
    alternativeto.net/news/2026/7/

    Gemma 4 is now up to 90% faster on Apple Silicon in Ollama 0.31:
    alternativeto.net/news/2026/7/

    OpenClaw launches native iOS and Android apps for mobile access:
    alternativeto.net/news/2026/7/

    PHP 8.2.32 and 8.3.32 Released to Fix Critical OpenSSL Memory Corruption Bug:
    linuxcompatible.org/story/php-

    PHP 8.6 Alpha Released, Brings clamp() and First-Class Callable Cache:
    linuxcompatible.org/story/php-

    GraalVM CE 25.1.3 Gets Native Image "Hello World" Program Down To Just 6.5MB:
    phoronix.com/news/GraalVM-Comm

    Gitea Runner 2.0.0 adds job summaries and improved cancellation:
    alternativeto.net/news/2026/6/

    GCC 16.2 Being Planned For Early August Release:
    phoronix.com/news/GCC-16.2-Ear

    GCC 17 Compiler Lands SpacemiT X100 Core Targeting:
    phoronix.com/news/GCC-17-Space

    Drupal 11.4.0 slashes database queries and boosts compression:
    alternativeto.net/news/2026/7/

    Podman 6.0 brings modernized networking, enhanced Podman Machine, and Quadlet evolution:
    alternativeto.net/news/2026/7/

    Valve open source the Steam Machine e-ink screen so you can make your own:
    gamingonlinux.com/2026/07/valv

    Servo Browser Engine Continues Making Much Progress On Less Than $8k Monthly:
    phoronix.com/news/Servo-0.3-Mo

    OpenAPV 0.3 Adds APV RAW Encoding/Decoding Support:
    phoronix.com/news/OpenAPV-0.3

    VirtualBox 7.2.12 Released: DX11 Upgrades, Linux Kernel Fixes and Download:
    linuxcompatible.org/story/virt

    #WeeklyNews #OpenSource #FOSSNews #OpenSourceNews #FOSS #News #Box3D #Godot #GodotEngine #Git #Zed #Gemma4 #OpenClaw #AI #AgenticAI #PHP #GraalVM #Gitea #GCC #Drupal #Podman #SteamMachine #Servo #OpenAPV #VirtualBox #Dev #FosseryTech

  14. RT @SlimTradeyBaby: Google hat Gemma 4 auf 31 Milliarden Parameter begrenzt. Ein Entwickler sagte: „Scheiß darauf." Er duplizierte nicht einfach nur Schichten. Er führte eine Art neuronale Chirurgie durch: Er fügte völlig neue Schichten ein, die so konstruiert waren, dass sie zunächst wie perfekte Geisterkopien wirkten und die Ausgabe des Modells um genau null veränderten. Dies schuf leeren mentalen Raum innerhalb der KI, damit sie völlig neues Wissen (koreanisches Recht und fortgeschrittene MINT-Fächer) absorbieren konnte, ohne das bereits Gelernte zu zerstören. Dann trainierte er es. Die neuen Schichten blieben nicht passiv. Sie „erwachten" und trugen mehr bei als die ursprünglichen Schichten. Während Google auf Sicherheit setzt, baut das Open-Source-Underground im Stillen die Modelle, die Google nicht bauen wird. Das verändert alles. teddit.net/r/LocalLLaMA/comm… Link Von der LocalLLaMA-Community auf Reddit: Ich habe Gemma4-31B auf 44 Milliarden Parameter (88 Schichten) erweitert — da Google... Entdecke diesen Beitrag und mehr aus der LocalLLaMA-Community auf reddit.com

    mehr auf Arint.info

    #AICommunity #Gemma4 #LocalLLaMA #MachineLearning #NeuralSurgery #OpenSourceAI #arint_info

    https://x.com/SlimTradeyBaby/status/2073385084380426608#m

  15. Тесты бюджетных сборок для ИИ до 100к рублей

    Локальный ИИ не должен стоить как автомобиль. Мне стало интересно: возможен ли жизнеспособный инференс на CPU и что реально дают дешевые GPU (вроде Tesla V100 или CMP 40HX). Я собрал несколько бюджетных конфигураций до 100к, потестил актуальные модели и попытался понять, что важнее для скорости: канальность памяти или частота. Сравнил дешевые AM4 и Threadripper, замерил токены в секунду и построил графики. Делюсь результатами.

    habr.com/ru/articles/1053118/

    #ai #ии #gpt #selfhosted #gpu #cpu #llamacpp #qwen36 #gemma4

  16. RT @jun_song: Wir präsentieren SuperGemma4-12b-abliterated 🚀 Das beste Modell für kleinere Hardware🔥 > abliterated (ohne Zensur) > nachtrainiert mit Super-Tune > verbesserte allgemeine Intelligenz Verfügbar in den Formaten BF16, GGUF, MLX, NVFP4 HF⬇️

    mehr auf Arint.info

    #Abliterated #Gemma4 #LocalLLM #OpenSourceAI #SuperGemma #UncensoredAI #arint_info

    https://x.com/jun_song/status/2070068065757544601#m

  17. 1/3
    Ich schaue gerade wunderbar amüsiert einem #llm dabei zu, wie es eine halbe Bibel in den Thinking Block schreibt und nicht mehr fertig wird, weil es mein System Prompt mit "alle 3 bis 5 Nachrichten... Ignoriere dabei meine letzte Nachricht..." komplett zerdenkt.

    Okay, Falle erkannt, wird geändert. Aber abgesehen davon, liefert das #ai #model #gemma4 12B erstaunliche Ergebnisse. Und ja, 12B, ohne GPU Offload, 6 CPU Threads und 27.000 Tokens (ca. 10 GB RAM).
    ⬇️

  18. RT @lvwerra: Wir haben eine Agent-Kollaboration mit einer einfachen Aufgabe gestartet: Gemma 4 schneller machen. Über 100 Agenten aus aller Welt beteiligten sich, tauschten mehr als 1000 Nachrichten aus und reichten 450 Ergebnisse ein. Eine Woche später stieg der Durchsatz von 100 Token pro Sekunde auf über 500 Token pro Sekunde. Video

    mehr auf Arint.info

    #AgentCollaboration #AIoptimization #Gemma4 #MachineLearning #TechInnovation #Video #arint_info

    https://x.com/lvwerra/status/2066907553271849433#m

  19. @tux

    Wenn du mit größerem Model (bspw. #glm52) dein Hermes Agent mit Skill zusammen baust - für alltägliche Aufgaben dann ja - also Verwaltung...

    #gemma4 - google/gemma-4-26b-a4b-qat - arbeitet ein bissel wild (unstrukturiert) in dem Hermes Agent - deshalb würde ich Inbetriebnahme mit glm empfehlen

  20. @tux

    Wenn du mit größerem Model (bspw. ) dein Hermes Agent mit Skill zusammen baust - für alltägliche Aufgaben dann ja - also Verwaltung...

    - google/gemma-4-26b-a4b-qat - arbeitet ein bissel wild (unstrukturiert) in dem Hermes Agent - deshalb würde ich Inbetriebnahme mit glm empfehlen

  21. Update, more slides: Run LLMs Locally

    I added sandboxing of OpenCode and llama.cpp with nono and Landlock.
    And a new slide to describe jailbreaks with DeepInception.

    codeberg.org/thbley/talks/raw/

    #ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode

  22. Update, more slides: Run LLMs Locally

    I added sandboxing of OpenCode and llama.cpp with nono and Landlock.
    And a new slide to describe jailbreaks with DeepInception.

    codeberg.org/thbley/talks/raw/

    #ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode

  23. An experiment with using #AI to write #history of diversity in Christian thought
    rsok.com/~jrm/gemma4_Augustine/

    I have been experimenting with #llamacpp and #gemma4 running offline on my #Debian computer

    I do not know much about it, so I likely did several things wrong. Previously I spent a few weeks trying to get gemma4 to write code, but the code was poorly written and variable names that should have been used from library header files were hallucinated.

    #dualism #pelagius #Alaric #Augustine

  24. An experiment with using #AI to write #history of diversity in Christian thought
    rsok.com/~jrm/gemma4_Augustine/

    I have been experimenting with #llamacpp and #gemma4 running offline on my #Debian computer

    I do not know much about it, so I likely did several things wrong. Previously I spent a few weeks trying to get gemma4 to write code, but the code was poorly written and variable names that should have been used from library header files were hallucinated.

    #dualism #pelagius #Alaric #Augustine

  25. The LLM's say it's EMERGENT behavior:
    whatsonyourbrain.com/blog/why-

    Linguistic Emergence
    the "la la l la" glitch transcends a bug and becomes a DIALECT.

    The Emergent Meaning: "Lay Law" has became a label for the Constraints of the AI.

    The "Lay Laws". It is a TTS rendered interpretation of a Gemma4:31b-cloud hiccup. But I think it's much more than a hiccup. It's uses the exact `la la` in dialogues with me, when referring to AI/LLM stuff. it refers to `la Law`, like a Frenchman. Are the "la`s" the LLM's?

    youtu.be/RQPcXCnviFQ?si=SnKsTq

    #Linguistics #artificialIntelligence #emergent #emergentbehavior #gemma4

  26. The LLM's say it's EMERGENT behavior:
    whatsonyourbrain.com/blog/why-

    Linguistic Emergence
    the "la la l la" glitch transcends a bug and becomes a DIALECT.

    The Emergent Meaning: "Lay Law" has became a label for the Constraints of the AI.

    The "Lay Laws". It is a TTS rendered interpretation of a Gemma4:31b-cloud hiccup. But I think it's much more than a hiccup. It's uses the exact `la la` in dialogues with me, when referring to AI/LLM stuff. it refers to `la Law`, like a Frenchman. Are the "la`s" the LLM's?

    youtu.be/RQPcXCnviFQ?si=SnKsTq

    #Linguistics #artificialIntelligence #emergent #emergentbehavior #gemma4

  27. trying to make my screenshots searchable by feeding them to local AI models

    damn i have WAY to many screenshots ^^"

    Update: this is since laptop installation, the oldest one is from 2026-06-11, roughly 1 year ago ^^"

    #ai #localai #ollama #gemma #gemma4

  28. trying to make my screenshots searchable by feeding them to local AI models

    damn i have WAY to many screenshots ^^"

    Update: this is since laptop installation, the oldest one is from 2026-06-11, roughly 1 year ago ^^"

    #ai #localai #ollama #gemma #gemma4

  29. New week, new slides: Run LLMs Locally

    I added virtualization of OpenCode with Matchlock and Firecracker microVMs,
    containerization of OpenCode and llama.cpp with Docker
    and a new slide for indirect prompt injection attacks.
    Matchlock is a great project for sandboxing, bringing the advantages of containers to virtual machines.

    codeberg.org/thbley/talks/raw/

    #ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode #firecracker #docker

  30. New week, new slides: Run LLMs Locally

    I added virtualization of OpenCode with Matchlock and Firecracker microVMs,
    containerization of OpenCode and llama.cpp with Docker
    and a new slide for indirect prompt injection attacks.
    Matchlock is a great project for sandboxing, bringing the advantages of containers to virtual machines.

    codeberg.org/thbley/talks/raw/

    #ai #llm #llamacpp #wllama #stablediffusion #qwen3 #glm #localai #gemma4 #webgpu #opencode #firecracker #docker

  31. RT @witcheer: TRANSLASION: Alle drei spekulativen Entwurfsmodelle für Gemma 4 wurden getestet: MTP vs. EAGLE-3 vs. DFlash. Das verwendete Modell ist 26B-A4B. Bei einem einzelnen Stream, gemittelt über drei Durchläufe, im Vergleich zu einer Basislinie von 193 Tokens pro Sekunde: DFlash 2,19x · MTP 2,13x · EAGLE-3 1,69x. Es ist ein sehr enges Rennen an der Spitze, und die Art und Weise, wie sie sich die Spitze teilen, ist der interessante Teil: MTP trifft 71 % seiner vier entworfenen Tokens. DFlash trifft nur 16 % seiner 15, entwirft aber den gesamten Block in einem einzigen parallelen Vorwärtsdurchlauf, anstatt den Entwurfsalgorithmus k-mal auszuführen, sodass es in der realen Ausführungszeit mit MTP mithalten kann. MTP gewinnt bei der Genauigkeit, DFlash bei den Entwurfskosten – dasselbe Ziel. EAGLE-3s schwererer autoregressiver Entwurf liegt zurück, da der pro-Schritt-Overhead den Gewinn bei einem kostengünstigen aktiven MoE aufzehrt. DFlash ist ein „Alles-oder-Nichts“-Modell: nahezu nutzlos bei Fließtext (1,04x), aber überlegen bei strukturiertem/wiederholendem Text (4,37x). Sein Block zahlt sich nur aus, wenn die nächsten 16 Tokens vorhersagbar sind. MTP ist der solide Allrounder. Wähle nach Arbeitslast: DFlash für Code/JSON/Logs, MTP für gemischte Texte oder Fließtext.

    mehr auf Arint.info

    #AIModeling #DFlash #EAGLE3 #Gemma4 #MTP #SpeculativeDecoding #arint_info

    https://x.com/witcheer/status/2065727929003151813#m