#ollama — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #ollama, aggregated by home.social.
-
Локальный ассистент для зумов, часть 2: как граф встреч становится памятью и почему ему можно верить
В первой части я собрал локального ассистента для созвонов: диаризация, стенограмма, граф знаний в Obsidian, ноль облаков. За полтора месяца ежедневной работы оказалось, что записать встречу — меньшая половина дела. Вторая часть — про то, как граф живёт неделями и почему ему можно верить: факты с датами «с» и «по», провенанс цитат из стенограммы, досье как этаж между поиском и графом, ночной прогон, песочница для облака, поиск по блокам и скрипт «забыть встречу целиком». Кода мало, устройства много. Читать дальше
https://habr.com/ru/articles/1078828/
#локальные_llm #ассистент_встреч #ollama #диаризация #граф_знаний #graphrag #obsidian #speechtotext #zoom #приватность
-
Локальный ассистент для зумов, часть 2: как граф встреч становится памятью и почему ему можно верить
В первой части я собрал локального ассистента для созвонов: диаризация, стенограмма, граф знаний в Obsidian, ноль облаков. За полтора месяца ежедневной работы оказалось, что записать встречу — меньшая половина дела. Вторая часть — про то, как граф живёт неделями и почему ему можно верить: факты с датами «с» и «по», провенанс цитат из стенограммы, досье как этаж между поиском и графом, ночной прогон, песочница для облака, поиск по блокам и скрипт «забыть встречу целиком». Кода мало, устройства много. Читать дальше
https://habr.com/ru/articles/1078828/
#локальные_llm #ассистент_встреч #ollama #диаризация #граф_знаний #graphrag #obsidian #speechtotext #zoom #приватность
-
Локальный ассистент для зумов, часть 2: как граф встреч становится памятью и почему ему можно верить
В первой части я собрал локального ассистента для созвонов: диаризация, стенограмма, граф знаний в Obsidian, ноль облаков. За полтора месяца ежедневной работы оказалось, что записать встречу — меньшая половина дела. Вторая часть — про то, как граф живёт неделями и почему ему можно верить: факты с датами «с» и «по», провенанс цитат из стенограммы, досье как этаж между поиском и графом, ночной прогон, песочница для облака, поиск по блокам и скрипт «забыть встречу целиком». Кода мало, устройства много. Читать дальше
https://habr.com/ru/articles/1078828/
#локальные_llm #ассистент_встреч #ollama #диаризация #граф_знаний #graphrag #obsidian #speechtotext #zoom #приватность
-
https://www.europesays.com/pl/668928/ NVIDIA PAIR – darmowe narzędzie rozdzielające zadania agentów AI między komputery w domowej sieci #AgenciAI #AiPc #DgxSpark #GPU #IFA2026 #LMStudio #LokalnaInferencja #Nauka #NaukaITechnika #NaukaTechnika #nvidia #ollama #PAIR #PersonalAIRouter #PL #Poland #Polish #Polska #Polski #RtxSpark #Science #ScienceAndTechnology #ScienceTechnology #SztucznaInteligencja #Technika #Technology
-
Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
@ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: https://javapro.io/2026/06/03/the-gen-ai-iceberg-java-tooling-edition/ -
Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
@ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: https://javapro.io/2026/06/03/the-gen-ai-iceberg-java-tooling-edition/ -
Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
@ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: https://javapro.io/2026/06/03/the-gen-ai-iceberg-java-tooling-edition/ -
Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
@ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: https://javapro.io/2026/06/03/the-gen-ai-iceberg-java-tooling-edition/ -
Most developers still think #Java + #AI means calling APIs from Spring Boot. The ecosystem goes deeper—from RAG frameworks to local inference with #Ollama or GPU workloads on #JVM.
@ArturSkowronski maps the #GenAI tooling iceberg for modern #Java: https://javapro.io/2026/06/03/the-gen-ai-iceberg-java-tooling-edition/ -
https://www.europesays.com/ie/672544/ Nvidia & Microsoft push local AI agents on Windows #acer #AIAgents(AgenticAI) #AIPC #ContentCreation #Creators #CyberLink #DataPrivacy #DeveloperTools #EdgeAI #EdgeComputing #Éire #ElectronicArts(EA) #GenerativeAI(GenAI) #HybridCloud #IE #Ireland #Lenovo #Linux #MacOS #Microsoft #MicrosoftWindows #Nvidia #Ollama #OpenSource #Optimisation #optimization #PCHardware #PCMarket #PortableComputer #RayTracing(RTX) #SmallOffice #Technology #Ubisoft
-
أطلق Raycast تحديثه v2، الذي يعيد دعم "Bring Your Own Model" (BYOM) لربط مزودين متوافقين مع OpenAI، أو تشغيل نماذج محلية عبر Ollama، أو استخدام OpenRouter API ضمن Raycast AI. هذه الميزات تتطلب الآن اشتراك Raycast Pro. كما يوسع التحديث إدارة النوافذ بأوامر دقيقة لتغيير الحجم والتحريك، وإمكانية إنشاء تخطيطات مخصصة من النوافذ الحالية. ويشمل التحسينات دعم النص من اليمين إلى اليسار في Raycast AI، وتحسينات في Quick AI، وتعديلات على ترتيب البحث عن الملفات.
-
أطلق Raycast تحديثه v2، الذي يعيد دعم "Bring Your Own Model" (BYOM) لربط مزودين متوافقين مع OpenAI، أو تشغيل نماذج محلية عبر Ollama، أو استخدام OpenRouter API ضمن Raycast AI. هذه الميزات تتطلب الآن اشتراك Raycast Pro. كما يوسع التحديث إدارة النوافذ بأوامر دقيقة لتغيير الحجم والتحريك، وإمكانية إنشاء تخطيطات مخصصة من النوافذ الحالية. ويشمل التحسينات دعم النص من اليمين إلى اليسار في Raycast AI، وتحسينات في Quick AI، وتعديلات على ترتيب البحث عن الملفات.
-
A short recap of my recent experiments of running an LLM.
Would be great to hear suggestions from others.
- Where else can we get decent hardware?
- What models do you run for your teams?
- How to reduce complexity of the setup?https://www.linkedin.com/pulse/recap-my-journey-self-host-llm-artur-neumann-6ddfe/
-
A short recap of my recent experiments of running an LLM.
Would be great to hear suggestions from others.
- Where else can we get decent hardware?
- What models do you run for your teams?
- How to reduce complexity of the setup?https://www.linkedin.com/pulse/recap-my-journey-self-host-llm-artur-neumann-6ddfe/
-
A short recap of my recent experiments of running an LLM.
Would be great to hear suggestions from others.
- Where else can we get decent hardware?
- What models do you run for your teams?
- How to reduce complexity of the setup?https://www.linkedin.com/pulse/recap-my-journey-self-host-llm-artur-neumann-6ddfe/
-
A short recap of my recent experiments of running an LLM.
Would be great to hear suggestions from others.
- Where else can we get decent hardware?
- What models do you run for your teams?
- How to reduce complexity of the setup?https://www.linkedin.com/pulse/recap-my-journey-self-host-llm-artur-neumann-6ddfe/
-
A short recap of my recent experiments of running an LLM.
Would be great to hear suggestions from others.
- Where else can we get decent hardware?
- What models do you run for your teams?
- How to reduce complexity of the setup?https://www.linkedin.com/pulse/recap-my-journey-self-host-llm-artur-neumann-6ddfe/
-
[Перевод] Qwen3.8-27B: лучший локальный LLM, который вы, вероятно, не сможете запустить
В этом месяце Alibaba выпустила две модели, и та, о которой все писали, оказалась не той, что мы ждали. Qwen3.8-Max — это API с 2,4 триллионами параметров, и пару недель назад я с большим энтузиазмом написал о ней обзор. Но релиз, которого я действительно ждал, вышел 14 августа: Qwen3.8-27B, Apache 2.0, веса на Hugging Face, модель достаточно компактна, чтобы работать на ноутбуке. Я должен кое в чем признаться. Я не провел полный тест этой модели. Я попробовал, на своем MacBook Air M4 получал около восьми токенов в секунду и потерял терпение где-то на втором запросе. По большей части этот пост посвящен именно моей неудаче, потому что я подозреваю, что у многих из вас вечер сложится так же, как у меня. Что представляет собой Qwen3.8-27B на самом деле 27,78 миллиарда параметров, плотная модель, включающая в себя визуальный энкодер, о котором никто не объявлял заранее. Принимает на вход текст, изображения и видео. Собственный контекст — 262 144 токена, с помощью YaRN можно увеличить его примерно до миллиона, если запускать модель на сервере. Архитектура — это по-настоящему интересная часть, и это не обычный трансформер. На протяжении 64 слоёв Qwen чередует 48 слоёв Gated DeltaNet (линейное внимание) с 16 полными слоями Gated Attention в соотношении 3:1. Только эти 16 слоёв с полным вниманием имеют KV-кэш. Таким образом, на каждый токен приходится около 64 КБ, что составляет примерно четверть от объёма памяти, необходимого для обычной 64-слойной модели с плотным кодированием. Если вы запомните только одно число из этого поста, пусть это будет именно оно. Всё, что касается того, поместится ли эта модель на вашем компьютере, зависит от объёма KV-кэша.
-
Hab mir im Februar ein #MacStudio mit #M4max und 128GB RAM gekauft... Für verrückte fast 4100€... Und ich dachte mir das ist der größte finanzielle Fehler ever
Hab damit dann viel #ollama #lmstudio #omlx gehostet - #LLM Experimentarium
Habs sogar ein paar Monaten nur mit den eigenen LLMs gecodet, aber Qualität naja...
Da sich die letzten Monate das #openWeights Business geändert hat und keine 120b Modelle veröffentlicht werden, die der perfekte fit dafür wären, nutze ich die 30b Modelle...
Leider mit zunehmendem Frust... #Gemma und #Qwen sind dafür einfach zu klein...
Dann bringt qwen auch nach 3.6 keine kleinen Modelle mehr raus - es gab kein 3.7 und kein 3.8 - letztes haben sie wieder man angekündigt
Bedeutet der Mac Studio war eigentlich nur als Desktop PC im Einsatz - eher #gaming und #Roblox programmierworkstation
Dann hab ich mich mal schlau gemacht - #apple hat meine 128gb Version zum frontier Modell gemacht - und die Modelle sind nirgendwo zu bekommen... Ganz merkwürdig Markt leer gesaugt
Vllt bringt Apple bald neue Mac Studio Modell raus???Okay - eBay preise abgecheckt - da gehen Modell für über 5k€ über den Tisch, aber sehr selten die 128gb Variante
eBay inseriert - 6999,99€ sofort kauf oder Preisvorschlag ab 6000€
Hab das Ding jetzt für 6200€ weiter verkauft
Leider über eBay bezahlt, deshalb wird mein Gewinn versteuert, aber ich gönne
Ist das nicht krass? Was da abgeht, das ist so verrückt auf dem Gebrauch Mac Markt!
-
От стримов к вебсокетам: как я боролся с буферизацией и наконец победил
Привет. Меня зовут Николай Пискунов, я руководитель направления Big Data и эксперт курса Cloud DevSecOps по безопасной разработке от Академии вАЙТИ
https://habr.com/ru/companies/beeline_cloud/articles/1057902/
#spring_ai #spring_boot #java #websocket #stomp #serversent_events #sse #ollama #llm #typescript
-
От стримов к вебсокетам: как я боролся с буферизацией и наконец победил
Привет. Меня зовут Николай Пискунов, я руководитель направления Big Data и эксперт курса Cloud DevSecOps по безопасной разработке от Академии вАЙТИ
https://habr.com/ru/companies/beeline_cloud/articles/1057902/
#spring_ai #spring_boot #java #websocket #stomp #serversent_events #sse #ollama #llm #typescript
-
От стримов к вебсокетам: как я боролся с буферизацией и наконец победил
Привет. Меня зовут Николай Пискунов, я руководитель направления Big Data и эксперт курса Cloud DevSecOps по безопасной разработке от Академии вАЙТИ
https://habr.com/ru/companies/beeline_cloud/articles/1057902/
#spring_ai #spring_boot #java #websocket #stomp #serversent_events #sse #ollama #llm #typescript
-
Confused by the exploding number of #AI tools in the #JVM ecosystem? Teams mix #SpringAI, #LangChain4j, MCP & #Ollama without understanding the layers underneath. Artur Skowronski explains what each part of the #Java AI stack is actually for: https://javapro.io/2026/06/03/the-gen-ai-iceberg-java-tooling-edition/
@langchain4j
-
Confused by the exploding number of #AI tools in the #JVM ecosystem? Teams mix #SpringAI, #LangChain4j, MCP & #Ollama without understanding the layers underneath. Artur Skowronski explains what each part of the #Java AI stack is actually for: https://javapro.io/2026/06/03/the-gen-ai-iceberg-java-tooling-edition/
@langchain4j
-
Confused by the exploding number of #AI tools in the #JVM ecosystem? Teams mix #SpringAI, #LangChain4j, MCP & #Ollama without understanding the layers underneath. Artur Skowronski explains what each part of the #Java AI stack is actually for: https://javapro.io/2026/06/03/the-gen-ai-iceberg-java-tooling-edition/
@langchain4j
-
Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.
I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. https://www.the-main-thread.com/p/quarkus-langchain4j-chat-memory-budget #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry
-
Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.
I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. https://www.the-main-thread.com/p/quarkus-langchain4j-chat-memory-budget #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry
-
Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.
I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. https://www.the-main-thread.com/p/quarkus-langchain4j-chat-memory-budget #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry
-
Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.
I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. https://www.the-main-thread.com/p/quarkus-langchain4j-chat-memory-budget #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry
-
Chat memory gets fuzzy fast once the UI hides what LangChain4j is actually retaining.
I wrote a Quarkus tutorial that makes retained-memory pressure visible with `TokenWindowChatMemory`, Ollama request counts, a turn ledger, and OpenTelemetry attributes. The useful split is simple: your app-level eviction budget is not the model context limit. https://www.the-main-thread.com/p/quarkus-langchain4j-chat-memory-budget #Java #Quarkus #LangChain4j #Ollama #OpenTelemetry
-
Local AI gets risky when the first confident answer becomes the system answer.
I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. https://www.the-main-thread.com/p/quarkus-langchain4j-ollama-signals #Java #Quarkus #LangChain4j #Ollama
-
Local AI gets risky when the first confident answer becomes the system answer.
I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. https://www.the-main-thread.com/p/quarkus-langchain4j-ollama-signals #Java #Quarkus #LangChain4j #Ollama
-
Local AI gets risky when the first confident answer becomes the system answer.
I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. https://www.the-main-thread.com/p/quarkus-langchain4j-ollama-signals #Java #Quarkus #LangChain4j #Ollama
-
Local AI gets risky when the first confident answer becomes the system answer.
I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. https://www.the-main-thread.com/p/quarkus-langchain4j-ollama-signals #Java #Quarkus #LangChain4j #Ollama
-
Local AI gets risky when the first confident answer becomes the system answer.
I wrote a Quarkus tutorial that sends the same text to two Ollama models, uses Quarkus Signals to escalate only on disagreement, and keeps `UNCERTAIN` separate from `FAILED`. https://www.the-main-thread.com/p/quarkus-langchain4j-ollama-signals #Java #Quarkus #LangChain4j #Ollama
-
Локальные LLM на Arch Linux и как увеличить скорость генерации в 20 раз
Приветствую всех читателей Хабра, в этой статье я хочу поделиться своим опытом в запуске локальных LLM, протестировать работоспособность интересных моделей на своем железе, рассказать, как я увеличил скорость генерации на одной из нейросетей в 20 раз (я не преувеличиваю). Но об этом чуть позже, а начну я повествование с описания своего железа.
https://habr.com/ru/articles/1045898/
#arch_linux #llamacpp #ollama #qwen36 #gemma4 #github #huggingface #intel_arc_b580
-
Confused by the exploding number of #AI tools in the #JVM ecosystem? Teams mix #SpringAI, #LangChain4j, MCP & #Ollama without understanding the layers underneath. Artur Skowronski explains what each part of the #Java AI stack is actually for: https://javapro.io/2026/06/03/the-gen-ai-iceberg-java-tooling-edition/
@langchain4j
-
Cheap questions should not burn the same local model as real debugging work.
I wrote a Quarkus + LangChain4j tutorial that classifies prompts, routes them between two Ollama models, and keeps the decision observable with CDI events and tests. https://www.the-main-thread.com/p/quarkus-langchain4j-model-routing #Java #Quarkus #LangChain4j #Ollama