home.social

#vectordatabase — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #vectordatabase, aggregated by home.social.

fetched live
  1. Dense Vectors in Machine Learning: Optimization Techniques Last week I was reviewing a retrieval pipeline that looked clean on paper and behaved strangely in production. Continue reading on Medium »

    #machine-learning #semantic-search #vector-database

    Origin | Interest | Match
  2. Never Lose the Right Chunk: How Hybrid Search Improves Recall in RAG Systems Table of Contents Continue reading on Medium »

    #artificial-intelligence #vector-database #retrieval-augmented-gen #llm #machine-learning

    Origin | Interest | Match
  3. I'm now experimenting with Open WebUI and local bge-m3 (embeddings), bge-reranker and Gemma-4-26b (instruct). I'm slowly learning how to integrate and test RAG systems. It's not so easy. If the instruct model is too eager to please it's not so clear if it uses the provided sources at all. I test with a rather "outdated" instruct model, llama-3.1-9b. It doesn't have enough stored information to "decode" more than a re-hash of the prompt.

    My current test set includes Nancy Leveson "Engineering a Safer World", Robert Rosen "Essays on Life Itself" and Enrico Martino "Intuitionistic Proof Versus Classical Truth". There are lots of structural similarities between "Systems Safety" and Rosennean (M,R)-systems. The connection to Brouwer is more subtle.

    My takeaway on RAG: people who just consume such solutions will have a hard time spotting shortcomings.

    Edit: after a lot of tuning my homelab RAG system is now useful.
    #homelab #RAG #vectordatabase

  4. Introducing OmniVec: An Open-Source Embedding Platform for AI Apps on Azure Originally posted on https://devblogs.microsoft.com/cosmosdb/introducing-omnivec-an-open-source-embedding-platform-for-ai...

    #ai #openai #vectordatabase #azure

    Origin | Interest | Match
  5. What Is a Vector Database? Why Traditional Databases Aren’t Enough for AI — Part 15 When I first heard about vector databases, I assumed they were just another database trend. Then I realized t...

    #generative-ai-tools #machine-learning #ai #llm #vector-database

    Origin | Interest | Match
  6. If your vector DB needs to see your data to search it, you’re not building private AI you’re renting confidence. “Private AI” has become one of the most overused phrases in modern infrastru...

    #ai #vectordatabase #beginners #llm

    Origin | Interest | Match
  7. Why AI Startups Are Rebuilding Search The biggest challenge in AI isn’t generating answers it’s finding the right information to generate them from. Continue reading on Medium »

    #endee #artificial-intelligence #vector-database #machine-learning #search-engines

    Origin | Interest | Match
  8. Embeddings in Generative AI: The Hidden Technology That Makes AI Actually Useful Why semantic search, RAG, recommendations, and AI assistants depend more on embeddings than most engineers realize. ...

    #machine-learning #generative-ai-development #vector-database #artificial-intelligence #embedding

    Origin | Interest | Match
  9. Why Your Vector Database Is Overpriced: Lucene's 32x Compression and Serverless Economics Why Your Vector Database Is Overpriced: Lucene's 32x Compression and Serverless Economics In 2026, ...

    #lucene #vectordatabase #serverless

    Origin | Interest | Match
  10. Sovereign Synapse: The Local Brain A vault of 3,150 Markdown files is just a very organized digital attic. It’s a repository of every conversation, code snippet, and research rabbit hole I’ve n...

    #python #ai #vectordatabase #localfirst

    Origin | Interest | Match
  11. Zilliz - Provides a managed vector database service.

    Cossmology Profile: dub.sh/HxFCcoG

    Key People: Charles Xie, James Luan

    #VectorDatabase #OpenSource #OSS #COSS

  12. Databases for : Should you use a vector ? 🤔

    This article compares projects competing to handle modern workloads, including and . Discover which databases best meet today’s AI challenges: lpi.org/636x

    (Disclaimer: This post contains an AI-generated image.)

  13. Your RAG’s Secret Backdoor: Leaking Data Through Vector Databases
    This article exposes a vulnerability in Retrieval-Augmented Generation (RAG) systems, where misconfigured vector databases can lead to sensitive data leakage. By improperly securing these databases, attackers can gain access to internal documents such as HR policies and top-secret product roadmaps. The RAG system works by storing document chunks as embeddings in a special-purpose vector database and querying it to provide context for the LLM. The focus on securing the LLM while neglecting the vector database leaves it vulnerable to data exfiltration. The attacker can exploit weak access controls and clever retrieval attacks to gain access to sensitive data. Key lesson: Secure vector databases to prevent data breaches caused by RAG system vulnerabilities. #BugBounty #ArtificialIntelligence #DataLeak #Infosec #VectorDatabase

    infosecwriteups.com/your-rags-

  14. Did you know? Our pgedge-vectorizer tool (on GitHub: github.com/pgEdge/pgedge-vecto) automatically chunks text content and generates vector embeddings with the help of background workers.

    OpenAI, Voyage AI, and Ollama are supported as embedding providers, and a simple SQL interface allows you to enable vectorization on any table. (There’s even built-in views and functions for monitoring queue status.)

    #github #opensource #semanticsearch #vector #vectordatabase #openai #ollama #voyageai

  15. Retrieval-Augmented Generation (RAG) Tutorial: Architecture, Implementation, and Production Guide:
    glukhov.org/rag/
    #AI #LLM #RAG #Embeddings #Reranking #VectorDatabase

  16. PostgreSQL with DiskANN indexing now beats Pinecone by 28x on latency at 75% lower cost. AdwaitX analyzes how OpenAI scaled to 800M users and why developers consolidate AI workloads. Technical breakdown #AdwaitX #PostgreSQL #VectorDatabase #AI

    adwaitx.com/postgresql-ai-appl

  17. If you're located near Illinois, Shaun Thomas will be presenting on "The New Postgres AI Ecosystem" at the Illinois Prairie PostgreSQL User Group this February 18th at 5:30 PM CST. 🐘

    Come by the DRW and say hi: meetup.com/illinois-prairie-po

    #postgresql #postgres #ai #vectordatabase #pgvector #vectorization #aidev #illinois #chicago

  18. pgedge-vectorizer: #Postgres extension that automatically vectorizes document contents and keeps vector embeddings current when the underlying content changes.

    Unlike other solutions, no external services or third party pipelines are required. It's also 100% open source under the #PostgreSQL license. ✨

    Check it out on GitHub: 👉 github.com/pgEdge/pgedge-vecto

    #programming #vector #vectordatabase #vectorsearch #vectordb #ai #llm #aiengineering #aidev #dba

  19. Trợ lý AI như ChatGPT thường quên lịch sử sau mỗi phiên làm việc, khiến người dùng phải mô tả lại lỗi nhiều lần. Bài viết đề xuất giải pháp: thêm lớp "bộ nhớ liên tục" dùng vector storage giữa CLI và AI, tự động lưu/lấy giải pháp cho các lỗi lặp lại. Braves este các thách thức như vệ sinh dữ liệu, định dạng fix lệnh chuẩn và bảo mật. Thử nghiệm CLI Python dùng DeepSeek V3 cho thấy giảm chi phí token về 0 cho sự cố đã giải quyết.

    #AI #CLI #DevTools #VectorDatabase #Programming
    #TríTuệNhânTạo #L

  20. SurgeDB: Cơ sở dữ liệu vector nhúng, hiệu năng cao, chạy nhẹ trên thiết bị biên, laptop hay VPS nhỏ. Viết bằng Rust, không phụ thuộc ngoài, hỗ trợ SIMD, HNSW, lọc metadata và bền vững với WAL. Chỉ tốn ~39MB RAM cho 100k vectors (768-dim), độ trễ tìm kiếm 0.64ms. Khác biệt với SQLite-vec và LanceDB ở kiến trúc hybrid in-memory + nén mạnh (SQ8/Binary). Đang tìm cộng sự phát triển WASM, tối ưu SIMD, binding Python/Node. Mã nguồn mở MIT. #SurgeDB #VectorDatabase #Rust #EdgeAI #HNSW #SQLite #AI #Mach

  21. A hands-on comparison of vector databases for RAG chatbots, showing why filtering and hybrid search matter in real production systems. hackernoon.com/how-to-choose-t #vectordatabase

  22. 🚀 ArcadeDB v25.12.1 is here! ✅ Fixed critical vector quantization bug ✅ New filtered vector search support ✅ Improved SQL functions & transaction logic ✅ 60+ dependency updates Release notes: github.com/ArcadeData/a... #ArcadeDB #VectorDatabase #OpenSource

    Release 25.12.1 · ArcadeData/a...

  23. 🚀 ArcadeDB v25.12.1 is here! ✅ Fixed critical vector quantization bug ✅ New filtered vector search support ✅ Improved SQL functions & transaction logic ✅ 60+ dependency updates Release notes: github.com/ArcadeData/a... #ArcadeDB #VectorDatabase #OpenSource

    Release 25.12.1 · ArcadeData/a...

  24. SrvDB v0.2.0 ra mắt: cơ sở dữ liệu vector offline, nhúng, không cần cloud. Hỗ trợ các chế độ chỉ mục Flat, HNSW, IVF, PQ + chế độ AUTO tự chọn dựa trên RAM/dataset. Cung cấp tìm kiếm chính xác & lượng tử, benchmark P99 latency, recall, disk, ingest trên laptop. Thiết kế cho RAG local, Edge/IoT, hệ thống air‑gapped và dev muốn thử mà không phụ thuộc cloud. Mong nhận phản hồi, báo cáo thực tế & góp ý. #SrvDB #VectorDatabase #AI #Edge #Offline #CôngNghệ #CơSởDữLiệu #AIlocal

    reddit.com/

  25. EdgeVec v0.7.0 ra mắt: Tìm kiếm vector ngay trên trình duyệt, không cần máy chủ hay API. Hỗ trợ giảm 32x bộ nhớ với binary quantization, tăng tốc 8.75x nhờ SIMD, lưu trữ bền vững bằng IndexedDB và lọc kết quả theo metadata. Hoàn hảo cho RAG offline, tìm kiếm tài liệu/code riêng tư, an toàn. Tất cả thao tác diễn ra trực tiếp trên thiết bị, không dữ liệu nào bị gửi ra ngoài.
    #EdgeVec #VectorDatabase #LocalLLM #WebAssembly #PrivacyFirst #AI #RAG #OfflineAI #VectorSearch #MachineLearning #Cơ_sở_dữ

  26. Tôi vừa xây dựng 1 vector database viết sẵn bằng C++, API bằng Go hỗ trợ các thao tác cơ bản. Hiện đang dùng bruteforce search để cải thiện, sắp chuyển sang HNSW. Mời bạn góp ý, test thử nghiệm, nhắn tin trao đổi repo nhé! #VectorDB #C++ #LậpTrìnhGo #PhátTriểnMở #VectorSearch #EarlyAdopters #VectorDatabase #HNSW #DevCommunity #NhàLậpTrình

    reddit.com/r/opensource/commen

  27. Did you know about pgedge-vectorizer? It's an open-source PostgreSQL extension that you can use to create vector embeddings for documents and keep the vectors automatically updated when the underlying content changes - no external services or third-party pipelines required, beyond an embedding LLM. 🤖

    Check it out on GitHub: 🔗 github.com/pgEdge/pgedge-vecto

    #opensource #postgresql #postgres #vector #ai #llm #aiengineering #programming #vectordatabase #database #data #devops #aiops

  28. Trying to set up semantic search for your #PostgreSQL database? Learn how to do things like...

    👉 Chunk documents into semantically meaningful pieces
    👉 Embed vectors for each chunk enabling similarity search
    👉 Automatically synchronize when documents change
    👉 Have all running inside PostgreSQL with no external services

    ...all in "Building a #RAG Server with PostgreSQL - Part 2: Chunking and Embeddings": 🔗 pgedge.com/blog/building-a-rag

    #programming #devops #postgres #vector #vectordatabase #pgvector

  29. In the final post of the "Building a RAG Server with PostgreSQL" series, learn how to deploy the open-source pgEdge RAG Server for a simple HTTP API you can use to ask questions about your content, and get answers based on your own documentation.

    Check it out 👉 pgedge.com/blog/building-a-rag

    Experiment with the RAG Server on #GitHub: github.com/pgEdge/pgedge-rag-s

    Want to see more of this kind of content? Let us know 💬

    #programming #ai #llm #vectordatabase #pgvector #vector #postgresql #postgres #data

  30. 🚀 ArcadeDB v25.11.1 is live! We've integrated the JVector engine for high-performance vector search, critical SQL fixes, smarter indexing for embedded lists, and improved gRPC serialization. github.com/ArcadeData/a... #ArcadeDB #OpenSource #GraphDB #VectorDatabase #NoSQL

    Release 25.11.1 · ArcadeData/a...

  31. Building a chat application over your own data often involves RAG and a vector database. Vector databases store and allow efficient searching of data embeddings. #VectorDatabase #ChatWithData

  32. A recent study suggests that stressful topics reduce the effectiveness of LLMs, also I'm experimenting with vector databases

    blog.salvius.org/2025/03/the-r

    #LLM #vectordatabase #redis #softwareengineering

Share on Mastodon

Enter the server where you have an account.