#rag — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #rag, aggregated by home.social.
-
Строим RAG вручную: как это работает под капотом
Имея некоторый опыт в построении классических ML и CV-проектов, я решил разобраться в NLP (Natural Language Processing) и собрать свою RAG-систему без использования сторонних RAG-фреймворков (LangChain, LlamaIndex и т.п.). Моя цель - понять, как на самом деле работает RAG под капотом: от разбиения текста до генерации ответа. Проблема, которую решает RAG Готовые LLM модели имеют несколько недостатков:
-
Честный RAG eval set: как собрать первые 100-300 кейсов и не обмануть себя цифрой
В прошлой статье серии мы сжимали эмбеддинги и замеряли, насколько изменится выдача относительно полного fp32-поиска. Там это был правильный вопрос: ломает ли квантизация уже существующий retrieval . Но у такого замера есть неприятное слепое пятно. Можно аккуратно сохранить 99% выдачи исходного эмбеддера и все равно плохо находить нужные документы на своем домене. Внешний бенчмарк, даже сильный, этого не гарантирует. BEIR как раз был создан, чтобы показать, насколько результаты retrieval меняются между разными задачами и доменами; один усредненный скор там не заменяет проверку на собственных данных. Нужен свой eval set . И тут обычно возникает ложная развилка:
https://habr.com/ru/articles/1070534/
#rag #retrieval #evals #оценка_качества #LLM #эмбеддинги #поиск #MTEB #ragas #BEIR
-
Наши книги о LLM: состояние дел по готовящимся новинкам
Приветствуем, Хабр. Не секрет, что большие языковые модели и, в частности, трансформеры (GPT) серьезно повлияли на работу программиста, привели к автоматизации многих рутинных задач, значительно удешевили проверку концепций и эксперименты при разработке новых продуктов. Коренные изменения произошли не только в разработке, но и во взаимодействии пользователя с ботами, агентами, поисковыми системами. Промпт-инжиниринг буквально за полтора года превратился из искусства в ремесло, которое способен освоить и подросток. Мы хотим очертить ближайшие перспективы выхода книг из типографии и планы на обозримое будущее — на наших верфях и уже практически на стапелях готовится целый флот литературы, ориентированной на работу с искусственным интеллектом.
https://habr.com/ru/companies/bhv_publishing/articles/1070496/
-
A tool failure is obvious. Bad retrieval can be much quieter.
The agent finds a plausible document, produces a grounded answer—and is confidently wrong.
From The AI Agent Test Manual:https://amzn.eu/d/0blIkveq
#AIAgents #RAG #LLMEvaluation #SoftwareTesting
https://webdad.eu/2026/08/14/the-agent-was-confident-the-data-was-wrong/
-
Milvus 2.6 で「接着剤コード」を捨てる ④ Decay Ranker と Boosting
https://qiita.com/sphereSky/items/60de6805e8a0381ce3c8?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
Milvus 2.6 で「接着剤コード」を捨てる ② Lexical Highlighting
https://qiita.com/sphereSky/items/f7414061db02a3530cd8?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
Milvus 2.6 で「接着剤コード」を捨てる ① Embedding Function
https://qiita.com/sphereSky/items/cc43ddaf827a46001158?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
on-premise "token generator" can become financially beneficial, with some analyses showing a break-even in as little as three months compared to public cloud API costs
500m input/output a mo for average smb - dev labs with heavy agentic pipelines will have 10x that
the mkt is wide open for smb open source ai consulting/praxis low tco, quick roi #cloud costs can vary by 68x
#iops #rag pipelines #orchestration #harness #hybrid models
-
on-premise "token generator" can become financially beneficial, with some analyses showing a break-even in as little as three months compared to public cloud API costs
500m input/output a mo for average smb - dev labs with heavy agentic pipelines will have 10x that
the mkt is wide open for smb open source ai consulting/praxis low tco, quick roi #cloud costs can vary by 68x
#iops #rag pipelines #orchestration #harness #hybrid models
-
J'avance doucement sur “RAG in my pocket“, une application pour du RAG personnel sur #RaspberryPI 5 avec les #SLM de chez Pleias : écriture de l'application #streamlit en cours depuis le notebook #Jupyter #IA #CPU #RAG #AI #Pi
-
J'avance doucement sur “RAG in my pocket“, une application pour du RAG personnel sur #RaspberryPI 5 avec les #SLM de chez Pleias : écriture de l'application #streamlit en cours depuis le notebook #Jupyter #IA #CPU #RAG #AI #Pi
-
IBM Cloud and Together AI expand AI infrastructure with NVIDIA https://www.cloudcomputing-news.net/news/ibm-together-ai-cloud-deal-nvidia-infrastructure/?utm_source=dlvr.it&utm_medium=mastodon #Cloud #Automation #Data #BusinessStrategy #DigitalTransformation #RAG #DataPlatforms #DigitalTransformation
-
IBM Cloud and Together AI expand AI infrastructure with NVIDIA https://www.cloudcomputing-news.net/news/ibm-together-ai-cloud-deal-nvidia-infrastructure/?utm_source=dlvr.it&utm_medium=mastodon #Cloud #Automation #Data #BusinessStrategy #DigitalTransformation #RAG #DataPlatforms #DigitalTransformation
-
From the Leanpub Blog: The Leanpub Podcast 🎙 Feat. Ediz Najim, Author of Systems Thinking for Agentic AI: A Software Architect’s Guide to Building Reliable LLM and Agent Systems
#books #leanpublishing #selfpublishing #AgenticAI #SystemsThinking #SoftwareArchitecture #LLMEngineering #AIReliability #RAG #ModelContextProtocol #BackendDevelopment #AIObservability #Leanpub
-
From the Leanpub Blog: The Leanpub Podcast 🎙 Feat. Ediz Najim, Author of Systems Thinking for Agentic AI: A Software Architect’s Guide to Building Reliable LLM and Agent Systems
#books #leanpublishing #selfpublishing #AgenticAI #SystemsThinking #SoftwareArchitecture #LLMEngineering #AIReliability #RAG #ModelContextProtocol #BackendDevelopment #AIObservability #Leanpub
-
NEW! The Leanpub Podcast 🎙 Feat. Ediz Najim, Author of Systems Thinking for Agentic AI: A Software Architect’s Guide to Building Reliable LLM and Agent Systems
#books #leanpublishing #selfpublishing #AgenticAI #SystemsThinking #SoftwareArchitecture #LLMEngineering #AIReliability #RAG #ModelContextProtocol #BackendDevelopment #AIObservability #Leanpub
-
NEW! The Leanpub Podcast 🎙 Feat. Ediz Najim, Author of Systems Thinking for Agentic AI: A Software Architect’s Guide to Building Reliable LLM and Agent Systems
#books #leanpublishing #selfpublishing #AgenticAI #SystemsThinking #SoftwareArchitecture #LLMEngineering #AIReliability #RAG #ModelContextProtocol #BackendDevelopment #AIObservability #Leanpub
-
Generic AI models hallucinate when they don't know your business. Ksolves' RAG development services ground AI in your actual data — for accurate, trustworthy responses.
🚀 Custom knowledge base integration
⚙️ Vector search & retrieval optimization
🔒 Secure data handling
🛡️ Ongoing model tuningTurn AI guesswork into grounded, reliable answers.
🔗 https://www.ksolves.com/rag-development-services
#RAG #Ksolves #GenerativeAI -
Generic AI models hallucinate when they don't know your business. Ksolves' RAG development services ground AI in your actual data — for accurate, trustworthy responses.
🚀 Custom knowledge base integration
⚙️ Vector search & retrieval optimization
🔒 Secure data handling
🛡️ Ongoing model tuningTurn AI guesswork into grounded, reliable answers.
🔗 https://www.ksolves.com/rag-development-services
#RAG #Ksolves #GenerativeAI -
RE: https://mastodon.gal/@damian/117078979222575784
@damian a que vén o de #teresinha?
Busqueino no diccionario da #RAG pero non sae.
Pensei que podía ser un nome local da #mantis que tampouco é moi galego.
-
Your vector search cannot tell "which services use Redis" from "which services do NOT use Redis". Both embed to nearly the same point, because the model encodes the topic and not the logic.
Same blind spot hits exact identifiers like ERR_CONNECTION_REFUSED or JIRA-4521, multi-hop questions, and anything time scoped.
Not a bug in your index, it is what embeddings are. Hybrid search, metadata filters and reranking are the patches.
-
Your vector search cannot tell "which services use Redis" from "which services do NOT use Redis". Both embed to nearly the same point, because the model encodes the topic and not the logic.
Same blind spot hits exact identifiers like ERR_CONNECTION_REFUSED or JIRA-4521, multi-hop questions, and anything time scoped.
Not a bug in your index, it is what embeddings are. Hybrid search, metadata filters and reranking are the patches.
-
Browser Policy Manager: документация
К релизу Browser Policy Manager 0.9.5 подготовлен документационный портал: четыре руководства на шести языках, локальный поиск, контекстная помощь из интерфейса и фундамент для будущего безопасного RAG-помощника. Рассказываю, как он устроен и почему документация здесь — часть безопасной работы с политиками Firefox.
https://habr.com/ru/articles/1069842/
#browser_policy_manager #firefox_enterprise #policiesjson #управление_политиками_браузера #техническая_документация #dita #локализация #rag #MPL20
-
holy motherf... shit... 🤯
i thought this would never end.... 🙈7 #ai #coding #agents build my new #RAG 📖
with a temporal #graph + #cli + #mcp) 🕸️ 🕐what a journey 😓
- real build time: 11 days
- lines of code (core): 33k
- lines of code (tests): 89k
- number of tests: 1868 (in 182 files)does it work: YES 😁
did i check any line of code: of course not 😂 🖕is it public ? sorry, no. the intensive live testing phase begins today 🏁
let's call it: expensive learning..... 😏
🧵 👇
-
holy motherf... shit... 🤯
i thought this would never end.... 🙈7 #ai #coding #agents build my new #RAG 📖
with a temporal #graph + #cli + #mcp) 🕸️ 🕐what a journey 😓
- real build time: 11 days
- lines of code (core): 33k
- lines of code (tests): 89k
- number of tests: 1868 (in 182 files)does it work: YES 😁
did i check any line of code: of course not 😂 🖕is it public ? sorry, no. the intensive live testing phase begins today 🏁
let's call it: expensive learning..... 😏
🧵 👇
-
LLM truncation in KoAssistant/Ollama is causing my prompt to be cut off.
The Fix:
1️⃣ Move from large models (e.g., 26B) to mid-size (e.g., 12B) to free up VRAM.
2️⃣ Create a Modelfile with PARAMETER num_ctx 32768.
3️⃣ Rebuild: ollama create model-name -f Modelfile.This balances intelligence and context, letting your RAG prompts actually reach the model! 🚀 #Ollama #LLM #LocalAI #KoReader #OpenSource #RAG #LocalLLM
-
Struggling with LLM truncation in KoAssistant/Ollama? 🛠️ If your prompt is being cut off, you're likely hitting VRAM limits.
The Fix:
1️⃣ Move from large models (e.g., 26B) to mid-size (e.g., 12B) to free up VRAM.
2️⃣ Create a Modelfile with PARAMETER num_ctx 32768.
3️⃣ Rebuild: ollama create model-name -f Modelfile.This balances intelligence and context, letting your RAG prompts actually reach the model! 🚀 #Ollama #LLM #LocalAI #KoReader #OpenSource #RAG #LocalLLM
-
Wohin mit dem Grubenwasser aus dem Steinkohlebergbau? Bisher leitet die RAG es teils in den Rhein. Doch der packt es nicht mehr.#WDR #Politik #Landespolitik #NRW #RAG #Rhein #Grubenwasser #Trinkwasser
Rhein zu niedrig - RAG muss Grubenwasser zurückhalten -
Wohin mit dem Grubenwasser aus dem Steinkohlebergbau? Bisher leitet die RAG es teils in den Rhein. Doch der packt es nicht mehr.#WDR #Politik #Landespolitik #NRW #RAG #Rhein #Grubenwasser #Trinkwasser
Rhein zu niedrig - RAG muss Grubenwasser zurückhalten -
Помогаем детским ревматологам повышать эффективность терапии: ИИ-агент для дистанционного мониторинга
В России около 70 тысяч детей и подростков страдают от ревматических патологий. Это группа заболеваний, при которых иммунная система даёт сбой и атакует собственные ткани — чаще всего суставы. Примерно раз в год ребёнку положена госпитализация в федеральный центр, чтобы проверить состояние и убедиться, что конкретный тип терапии работает. Однако, когда ребёнок возвращается в родной город, рядом может не быть ни одного врача, который знает, что делать с этой терапией. На связи Юлия Шеянова, менеджер проектов в здравоохранении, Центр технологий для общества Yandex Cloud. В этой статье расскажу, как мы в Yandex Cloud вместе с Ассоциацией детских ревматологов (ASPiRRe) и студентом ИТМО сделали ИИ-агента для системы дистанционного мониторинга. Ниже — как он устроен и почему мы доверили модели только две узкие задачи, а все решения о рисках вынесли в детерминированные правила.
https://habr.com/ru/companies/yandex_cloud_and_infra/articles/1068692/
-
What Happens When a Code-Signing Key Is Stolen? https://www.cloudcomputing-news.net/news/what-happens-when-a-code-signing-key-is-stolen/?utm_source=dlvr.it&utm_medium=mastodon #Cloud #Automation #Data #RAG #BusinessStrategy #AIArchitecture #GenerativeAI #DataScience
-
What Happens When a Code-Signing Key Is Stolen? https://www.cloudcomputing-news.net/news/what-happens-when-a-code-signing-key-is-stolen/?utm_source=dlvr.it&utm_medium=mastodon #Cloud #Automation #Data #RAG #BusinessStrategy #AIArchitecture #GenerativeAI #DataScience
-
RAGに古い情報で答えさせないためには?"コンバージド"データベースから考えるAI時代のデータ基盤
https://qiita.com/yushibats/items/9dd91baaa89c919d0992?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items -
From the Leanpub Blog: Leanpub Book LAUNCH 🚀 LLM Engineering, from Component to Production by Ali Aouf
#books #leanpublishing #selfpublishing #LLMEngineering #RAG #AIAgents #GenerativeAI #MLOps
-
From the Leanpub Blog: Leanpub Book LAUNCH 🚀 LLM Engineering, from Component to Production by Ali Aouf
#books #leanpublishing #selfpublishing #LLMEngineering #RAG #AIAgents #GenerativeAI #MLOps
-
NEW! Leanpub Book LAUNCH 🚀 LLM Engineering, from Component to Production by Ali Aouf
#books #leanpublishing #selfpublishing #LLMEngineering #RAG #AIAgents #GenerativeAI #MLOps
-
NEW! Leanpub Book LAUNCH 🚀 LLM Engineering, from Component to Production by Ali Aouf
#books #leanpublishing #selfpublishing #LLMEngineering #RAG #AIAgents #GenerativeAI #MLOps
-
OCI Data Catalog 第3回:Object Storage上のOracle DatabaseマニュアルPDFでSelect AI with RAGを構築してみてみた
https://qiita.com/shirok/items/704f8f380fee205e06d1?utm_campaign=popular_items&utm_medium=feed&utm_source=popular_items#qiita #ObjectStorage #oraclecloud #rag #datacatalog #SelectAI
-
RT @thesupermanmx: Google DeepMind argues RAG is broken. They published a paper that proved vectors databases are the dead end. For the last three years, the default engineering response to any AI memory or data problem has been identical: "Just build a RAG pipeline." Chunk the data, push it into a vector database, and let embeddings handle the rest. Every company scaling enterprise AI assumes that if an embedding model fails, it's just a matter of time. Better training data, larger models, more parameters—throw compute at it, and the search gets smarter. This paper proves that assumption is completely false. They mathematically demonstrated that single-vector embeddings have a hard, uncrossable limit. Here is the core flaw: An embedding compresses an entire document or a complex query down into a single fixed-length vector of numbers. When you run a search, the model takes the dot product of those vectors to measure similarity. The math reveals a brutal constraint. The number of distinct document combinations a model can possibly retrieve for different queries is strictly bounded by the dimension of its embedding space. It is a hard mathematical ceiling dictated by geometry and communication complexity. No amount of data scaling can fix it. No amount of fine-tuning will punch through it. Even if you give an embedding model infinite, unconstrained training freedom on the test set, it still hits the wall. DeepMind built a stress-test dataset called LIMIT to prove it. They threw state-of-the-art embedding models at it, models with thousands of dimensions. The models completely failed. Even on s…
mehr auf Arint.info
#agent #AIagent #DeepMind #finetuning #Google #RAG #rest #arint_info
-
RT @thesupermanmx: Google DeepMind argues RAG is broken. They published a paper that proved vectors databases are the dead end. For the last three years, the default engineering response to any AI memory or data problem has been identical: "Just build a RAG pipeline." Chunk the data, push it into a vector database, and let embeddings handle the rest. Every company scaling enterprise AI assumes that if an embedding model fails, it's just a matter of time. Better training data, larger models, more parameters—throw compute at it, and the search gets smarter. This paper proves that assumption is completely false. They mathematically demonstrated that single-vector embeddings have a hard, uncrossable limit. Here is the core flaw: An embedding compresses an entire document or a complex query down into a single fixed-length vector of numbers. When you run a search, the model takes the dot product of those vectors to measure similarity. The math reveals a brutal constraint. The number of distinct document combinations a model can possibly retrieve for different queries is strictly bounded by the dimension of its embedding space. It is a hard mathematical ceiling dictated by geometry and communication complexity. No amount of data scaling can fix it. No amount of fine-tuning will punch through it. Even if you give an embedding model infinite, unconstrained training freedom on the test set, it still hits the wall. DeepMind built a stress-test dataset called LIMIT to prove it. They threw state-of-the-art embedding models at it, models with thousands of dimensions. The models completely failed. Even on s…
mehr auf Arint.info
#agent #AIagent #DeepMind #finetuning #Google #RAG #rest #arint_info
-
Я объяснил KERNEL маме за пять минут. Потом проверил, переживёт ли он LLM в production
Недавно я попробовал объяснить маме, как нормально ставить задачи ChatGPT. Без temperature, context window, system prompt и прочих слов, после которых человек вполне справедливо решает, что проще уже сделать всё самому. Я сказал примерно так: «Сформулируй, что тебе нужно. Объясни, каким должен быть результат. Если есть ограничения, напиши их сразу. Не сваливай пять разных задач в одну. И не заставляй модель угадывать то, что существует только у тебя в голове». На всё ушло минут пять. А потом я понял, что почти дословно пересказал KERNEL, один из фреймворков для написания промптов. И тут стало интереснее: если базовый prompt engineering действительно можно объяснить человеку за пять минут, что происходит, когда LLM переезжает из вкладки браузера в реальный продукт?
https://habr.com/ru/companies/syntx_ai/articles/1068654/
#LLM #prompt_engineering #KERNEL #промпты #context_engineering #evals #ChatGPT #RAG #AIагенты #LLM_в_production
-
From the Leanpub Blog: Leanpub Book LAUNCH 🚀 ENTERPRISE AI ARCHITECTURE AND THE MODERN AI STACK: VOLUME I — DESIGNING THE STACK by Padmanabham Venkiteela
#books #leanpublishing #selfpublishing #EnterpriseAI #AIArchitecture #LLM #AgenticAI #RAG
-
From the Leanpub Blog: Leanpub Book LAUNCH 🚀 ENTERPRISE AI ARCHITECTURE AND THE MODERN AI STACK: VOLUME I — DESIGNING THE STACK by Padmanabham Venkiteela
#books #leanpublishing #selfpublishing #EnterpriseAI #AIArchitecture #LLM #AgenticAI #RAG
-
NEW! Leanpub Book LAUNCH 🚀 ENTERPRISE AI ARCHITECTURE AND THE MODERN AI STACK: VOLUME I — DESIGNING THE STACK by Padmanabham Venkiteela
#books #leanpublishing #selfpublishing #EnterpriseAI #AIArchitecture #LLM #AgenticAI #RAG
-
NEW! Leanpub Book LAUNCH 🚀 ENTERPRISE AI ARCHITECTURE AND THE MODERN AI STACK: VOLUME I — DESIGNING THE STACK by Padmanabham Venkiteela
#books #leanpublishing #selfpublishing #EnterpriseAI #AIArchitecture #LLM #AgenticAI #RAG
-
Debugging a fine-tuned model is archaeology. Debugging RAG is grep.
Fine-tune fix: 3-5 days engineering, $500-5,000 compute, regression risk on every retrain. RAG fix: edit the source document, 30 minutes, zero compute.
60% of production LLM apps now run on RAG. Not because it is more sophisticated, but because keeping a database current is a solved problem.
roamingpigs.com/r/rvfte
-
What is Retrieval-Augmented Generation (RAG)? 🤖
RAG is an AI technique that combines information retrieval with generative AI to produce more contextually informed answers.
Our beginner-friendly guide covers:
• What RAG means
• How RAG works
• RAG architecture
• Benefits and limitations
• Real-world examplesRead the full guide:
https://startupglossary.blogspot.com/2026/08/what-is-rag.html -
Nouveau billet, et le début d'une série sur mes projets : des assistants IA pour réparer de vieilles machines — une Datsun de 77, un flipper Dracula, des archives techniques du XIXe.
Le fil rouge : une IA qui ne répond qu'à partir de sources réelles, cite ses pages, et dit « je ne sais pas » plutôt que d'inventer. La machine propose, l'humain dispose.
https://blogz.zaclys.com/chatpey/une-ia-qui-ne-ment-pas-sur-ce-quelle-ne-sait-pas