home.social

#ollama — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #ollama, aggregated by home.social.

  1. Локальный ассистент для зумов, часть 2: как граф встреч становится памятью и почему ему можно верить

    В первой части я собрал локального ассистента для созвонов: диаризация, стенограмма, граф знаний в Obsidian, ноль облаков. За полтора месяца ежедневной работы оказалось, что записать встречу — меньшая половина дела. Вторая часть — про то, как граф живёт неделями и почему ему можно верить: факты с датами «с» и «по», провенанс цитат из стенограммы, досье как этаж между поиском и графом, ночной прогон, песочница для облака, поиск по блокам и скрипт «забыть встречу целиком». Кода мало, устройства много. Читать дальше

    habr.com/ru/articles/1078828/

    #локальные_llm #ассистент_встреч #ollama #диаризация #граф_знаний #graphrag #obsidian #speechtotext #zoom #приватность

  2. Локальный ассистент для зумов, часть 2: как граф встреч становится памятью и почему ему можно верить

    В первой части я собрал локального ассистента для созвонов: диаризация, стенограмма, граф знаний в Obsidian, ноль облаков. За полтора месяца ежедневной работы оказалось, что записать встречу — меньшая половина дела. Вторая часть — про то, как граф живёт неделями и почему ему можно верить: факты с датами «с» и «по», провенанс цитат из стенограммы, досье как этаж между поиском и графом, ночной прогон, песочница для облака, поиск по блокам и скрипт «забыть встречу целиком». Кода мало, устройства много. Читать дальше

    habr.com/ru/articles/1078828/

    #локальные_llm #ассистент_встреч #ollama #диаризация #граф_знаний #graphrag #obsidian #speechtotext #zoom #приватность

  3. Локальный ассистент для зумов, часть 2: как граф встреч становится памятью и почему ему можно верить

    В первой части я собрал локального ассистента для созвонов: диаризация, стенограмма, граф знаний в Obsidian, ноль облаков. За полтора месяца ежедневной работы оказалось, что записать встречу — меньшая половина дела. Вторая часть — про то, как граф живёт неделями и почему ему можно верить: факты с датами «с» и «по», провенанс цитат из стенограммы, досье как этаж между поиском и графом, ночной прогон, песочница для облака, поиск по блокам и скрипт «забыть встречу целиком». Кода мало, устройства много. Читать дальше

    habr.com/ru/articles/1078828/

    #локальные_llm #ассистент_встреч #ollama #диаризация #граф_знаний #graphrag #obsidian #speechtotext #zoom #приватность

  4. @Madic noch darauf Bezug nehmen:

    Bei #ollama Cloud bekommst du in pro abo für 20$ echtes Geld 60$ token Guthaben... Also Faktor 3 - dann ist es ja dauerhaft 66% reduziert dort

  5. @Madic noch darauf Bezug nehmen:

    Bei #ollama Cloud bekommst du in pro abo für 20$ echtes Geld 60$ token Guthaben... Also Faktor 3 - dann ist es ja dauerhaft 66% reduziert dort

  6. @Madic noch darauf Bezug nehmen:

    Bei #ollama Cloud bekommst du in pro abo für 20$ echtes Geld 60$ token Guthaben... Also Faktor 3 - dann ist es ja dauerhaft 66% reduziert dort

  7. @Madic noch darauf Bezug nehmen:

    Bei Cloud bekommst du in pro abo für 20$ echtes Geld 60$ token Guthaben... Also Faktor 3 - dann ist es ja dauerhaft 66% reduziert dort

  8. Learn when to migrate from Ollama to vLLM. Migration signals, planning steps, Docker Compose setup, and a practical checklist for moving your local LLM server.

    #Ollama #vLLM #LLM #AI #Self-Hosting #Docker #API #DevOps

    glukhov.org/llm-hosting/compar

  9. Learn when to migrate from Ollama to vLLM. Migration signals, planning steps, Docker Compose setup, and a practical checklist for moving your local LLM server.

    #Ollama #vLLM #LLM #AI #Self-Hosting #Docker #API #DevOps

    glukhov.org/llm-hosting/compar

  10. Learn when to migrate from Ollama to vLLM. Migration signals, planning steps, Docker Compose setup, and a practical checklist for moving your local LLM server.

    #Ollama #vLLM #LLM #AI #Self-Hosting #Docker #API #DevOps

    glukhov.org/llm-hosting/compar

  11. Learn when to migrate from Ollama to vLLM. Migration signals, planning steps, Docker Compose setup, and a practical checklist for moving your local LLM server.

    #Ollama #vLLM #LLM #AI #Self-Hosting #Docker #API #DevOps

    glukhov.org/llm-hosting/compar

  12. Learn when to migrate from Ollama to vLLM. Migration signals, planning steps, Docker Compose setup, and a practical checklist for moving your local LLM server.

    -Hosting

    glukhov.org/llm-hosting/compar

  13. 📝 Daily report 📈

    Here are today's most popular trending hashtags #⃣ on our website 🌐️:

    #openai, #gnu, #sysadmin, #ollama

    🔥 Stay tuned! 🔥

  14. 📝 Daily report 📈

    Here are today's most popular trending hashtags #⃣ on our website 🌐️:

    #openai, #gnu, #sysadmin, #ollama

    🔥 Stay tuned! 🔥

  15. 📝 Daily report 📈

    Here are today's most popular trending hashtags #⃣ on our website 🌐️:

    #openai, #gnu, #sysadmin, #ollama

    🔥 Stay tuned! 🔥

  16. Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
    @ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: javapro.io/2026/06/03/the-gen-

    @ollama

  17. Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
    @ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: javapro.io/2026/06/03/the-gen-

    @ollama

  18. Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
    @ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: javapro.io/2026/06/03/the-gen-

    @ollama

  19. Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
    @ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: javapro.io/2026/06/03/the-gen-

    @ollama

  20. Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
    @ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: javapro.io/2026/06/03/the-gen-

    @ollama

  21. أطلق Raycast تحديثه v2، الذي يعيد دعم "Bring Your Own Model" (BYOM) لربط مزودين متوافقين مع OpenAI، أو تشغيل نماذج محلية عبر Ollama، أو استخدام OpenRouter API ضمن Raycast AI. هذه الميزات تتطلب الآن اشتراك Raycast Pro. كما يوسع التحديث إدارة النوافذ بأوامر دقيقة لتغيير الحجم والتحريك، وإمكانية إنشاء تخطيطات مخصصة من النوافذ الحالية. ويشمل التحسينات دعم النص من اليمين إلى اليسار في Raycast AI، وتحسينات في Quick AI، وتعديلات على ترتيب البحث عن الملفات.

    #Raycast #AI #Ollama

  22. أطلق Raycast تحديثه v2، الذي يعيد دعم "Bring Your Own Model" (BYOM) لربط مزودين متوافقين مع OpenAI، أو تشغيل نماذج محلية عبر Ollama، أو استخدام OpenRouter API ضمن Raycast AI. هذه الميزات تتطلب الآن اشتراك Raycast Pro. كما يوسع التحديث إدارة النوافذ بأوامر دقيقة لتغيير الحجم والتحريك، وإمكانية إنشاء تخطيطات مخصصة من النوافذ الحالية. ويشمل التحسينات دعم النص من اليمين إلى اليسار في Raycast AI، وتحسينات في Quick AI، وتعديلات على ترتيب البحث عن الملفات.

    #Raycast #AI #Ollama

  23. A short recap of my recent experiments of running an LLM.

    Would be great to hear suggestions from others.
    - Where else can we get decent hardware?
    - What models do you run for your teams?
    - How to reduce complexity of the setup?

    linkedin.com/pulse/recap-my-jo

    #AI #LLM #ollama #openwebui #hetzner #trooperai #selfhosted

  24. A short recap of my recent experiments of running an LLM.

    Would be great to hear suggestions from others.
    - Where else can we get decent hardware?
    - What models do you run for your teams?
    - How to reduce complexity of the setup?

    linkedin.com/pulse/recap-my-jo

    #AI #LLM #ollama #openwebui #hetzner #trooperai #selfhosted

  25. A short recap of my recent experiments of running an LLM.

    Would be great to hear suggestions from others.
    - Where else can we get decent hardware?
    - What models do you run for your teams?
    - How to reduce complexity of the setup?

    linkedin.com/pulse/recap-my-jo

    #AI #LLM #ollama #openwebui #hetzner #trooperai #selfhosted

  26. A short recap of my recent experiments of running an LLM.

    Would be great to hear suggestions from others.
    - Where else can we get decent hardware?
    - What models do you run for your teams?
    - How to reduce complexity of the setup?

    linkedin.com/pulse/recap-my-jo

    #AI #LLM #ollama #openwebui #hetzner #trooperai #selfhosted

  27. A short recap of my recent experiments of running an LLM.

    Would be great to hear suggestions from others.
    - Where else can we get decent hardware?
    - What models do you run for your teams?
    - How to reduce complexity of the setup?

    linkedin.com/pulse/recap-my-jo

    #AI #LLM #ollama #openwebui #hetzner #trooperai #selfhosted

  28. [Перевод] Qwen3.8-27B: лучший локальный LLM, который вы, вероятно, не сможете запустить

    В этом месяце Alibaba выпустила две модели, и та, о которой все писали, оказалась не той, что мы ждали. Qwen3.8-Max — это API с 2,4 триллионами параметров, и пару недель назад я с большим энтузиазмом написал о ней обзор. Но релиз, которого я действительно ждал, вышел 14 августа: Qwen3.8-27B, Apache 2.0, веса на Hugging Face, модель достаточно компактна, чтобы работать на ноутбуке. Я должен кое в чем признаться. Я не провел полный тест этой модели. Я попробовал, на своем MacBook Air M4 получал около восьми токенов в секунду и потерял терпение где-то на втором запросе. По большей части этот пост посвящен именно моей неудаче, потому что я подозреваю, что у многих из вас вечер сложится так же, как у меня. Что представляет собой Qwen3.8-27B на самом деле 27,78 миллиарда параметров, плотная модель, включающая в себя визуальный энкодер, о котором никто не объявлял заранее. Принимает на вход текст, изображения и видео. Собственный контекст — 262 144 токена, с помощью YaRN можно увеличить его примерно до миллиона, если запускать модель на сервере. Архитектура — это по-настоящему интересная часть, и это не обычный трансформер. На протяжении 64 слоёв Qwen чередует 48 слоёв Gated DeltaNet (линейное внимание) с 16 полными слоями Gated Attention в соотношении 3:1. Только эти 16 слоёв с полным вниманием имеют KV-кэш. Таким образом, на каждый токен приходится около 64 КБ, что составляет примерно четверть от объёма памяти, необходимого для обычной 64-слойной модели с плотным кодированием. Если вы запомните только одно число из этого поста, пусть это будет именно оно. Всё, что касается того, поместится ли эта модель на вашем компьютере, зависит от объёма KV-кэша.

    habr.com/ru/articles/1072048/

    #qwen #llm #quantization #apple_silicon #lmstudio #ollama

  29. Hab mir im Februar ein #MacStudio mit #M4max und 128GB RAM gekauft... Für verrückte fast 4100€... Und ich dachte mir das ist der größte finanzielle Fehler ever

    Hab damit dann viel #ollama #lmstudio #omlx gehostet - #LLM Experimentarium

    Habs sogar ein paar Monaten nur mit den eigenen LLMs gecodet, aber Qualität naja...

    Da sich die letzten Monate das #openWeights Business geändert hat und keine 120b Modelle veröffentlicht werden, die der perfekte fit dafür wären, nutze ich die 30b Modelle...

    Leider mit zunehmendem Frust... #Gemma und #Qwen sind dafür einfach zu klein...

    Dann bringt qwen auch nach 3.6 keine kleinen Modelle mehr raus - es gab kein 3.7 und kein 3.8 - letztes haben sie wieder man angekündigt

    Bedeutet der Mac Studio war eigentlich nur als Desktop PC im Einsatz - eher #gaming und #Roblox programmierworkstation

    Dann hab ich mich mal schlau gemacht - #apple hat meine 128gb Version zum frontier Modell gemacht - und die Modelle sind nirgendwo zu bekommen... Ganz merkwürdig Markt leer gesaugt
    Vllt bringt Apple bald neue Mac Studio Modell raus???

    Okay - eBay preise abgecheckt - da gehen Modell für über 5k€ über den Tisch, aber sehr selten die 128gb Variante

    eBay inseriert - 6999,99€ sofort kauf oder Preisvorschlag ab 6000€

    Hab das Ding jetzt für 6200€ weiter verkauft

    Leider über eBay bezahlt, deshalb wird mein Gewinn versteuert, aber ich gönne

    Ist das nicht krass? Was da abgeht, das ist so verrückt auf dem Gebrauch Mac Markt!

  30. От стримов к вебсокетам: как я боролся с буферизацией и наконец победил

    Привет. Меня зовут Николай Пискунов, я руководитель направления Big Data и эксперт курса Cloud DevSecOps по безопасной разработке от Академии вАЙТИ

    habr.com/ru/companies/beeline_

    #spring_ai #spring_boot #java #websocket #stomp #serversent_events #sse #ollama #llm #typescript

  31. От стримов к вебсокетам: как я боролся с буферизацией и наконец победил

    Привет. Меня зовут Николай Пискунов, я руководитель направления Big Data и эксперт курса Cloud DevSecOps по безопасной разработке от Академии вАЙТИ

    habr.com/ru/companies/beeline_

    #spring_ai #spring_boot #java #websocket #stomp #serversent_events #sse #ollama #llm #typescript

  32. От стримов к вебсокетам: как я боролся с буферизацией и наконец победил

    Привет. Меня зовут Николай Пискунов, я руководитель направления Big Data и эксперт курса Cloud DevSecOps по безопасной разработке от Академии вАЙТИ

    habr.com/ru/companies/beeline_

    #spring_ai #spring_boot #java #websocket #stomp #serversent_events #sse #ollama #llm #typescript

  33. Confused by the exploding number of #AI tools in the #JVM ecosystem? Teams mix #SpringAI, #LangChain4j, MCP & #Ollama without understanding the layers underneath. Artur Skowronski explains what each part of the #Java AI stack is actually for: javapro.io/2026/06/03/the-gen-

    @langchain4j

  34. Confused by the exploding number of #AI tools in the #JVM ecosystem? Teams mix #SpringAI, #LangChain4j, MCP & #Ollama without understanding the layers underneath. Artur Skowronski explains what each part of the #Java AI stack is actually for: javapro.io/2026/06/03/the-gen-

    @langchain4j

  35. Confused by the exploding number of #AI tools in the #JVM ecosystem? Teams mix #SpringAI, #LangChain4j, MCP & #Ollama without understanding the layers underneath. Artur Skowronski explains what each part of the #Java AI stack is actually for: javapro.io/2026/06/03/the-gen-

    @langchain4j

  36. Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.

    I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry

  37. Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.

    I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry

  38. Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.

    I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry

  39. Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.

    I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry

  40. Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.

    I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry

  41. Local AI gets risky when the first confident answer becomes the system answer.

    I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama

  42. Local AI gets risky when the first confident answer becomes the system answer.

    I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama

  43. Local AI gets risky when the first confident answer becomes the system answer.

    I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama

  44. Local AI gets risky when the first confident answer becomes the system answer.

    I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama

  45. Local AI gets risky when the first confident answer becomes the system answer.

    I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama

  46. Локальные LLM на Arch Linux и как увеличить скорость генерации в 20 раз

    Приветствую всех читателей Хабра, в этой статье я хочу поделиться своим опытом в запуске локальных LLM, протестировать работоспособность интересных моделей на своем железе, рассказать, как я увеличил скорость генерации на одной из нейросетей в 20 раз (я не преувеличиваю). Но об этом чуть позже, а начну я повествование с описания своего железа.

    habr.com/ru/articles/1045898/

    #arch_linux #llamacpp #ollama #qwen36 #gemma4 #github #huggingface #intel_arc_b580

  47. Confused by the exploding number of #AI tools in the #JVM ecosystem? Teams mix #SpringAI, #LangChain4j, MCP & #Ollama without understanding the layers underneath. Artur Skowronski explains what each part of the #Java AI stack is actually for: javapro.io/2026/06/03/the-gen-

    @langchain4j

  48. Cheap questions should not burn the same local model as real debugging work.

    I wrote a Quarkus + LangChain4j tutorial that classifies prompts, routes them between two Ollama models, and keeps the decision observable with CDI events and tests. the-main-thread.com/p/quarkus- #Java #Quarkus #LangChain4j #Ollama