RAG & Retrieval

Vector Databases for RAG: How to Choose One and Wire It In (2026)

The vector database is the one component of a RAG pipeline you can't skip — but you may not need a dedicated one. This guide explains what a vector database for RAG actually does, when pgvector inside Postgres is enough, how the main open-source and managed options compare (Chroma, Qdrant, Weaviate, Milvus, Pinecone), a minimal working example with Chroma, and how to check the retrieval it returns is any good using retrieval evaluation metrics.

Gurram Poorna Prudhvi

Lead AI Engineer

Technical Guide
Sep 30, 2026
11 min read
VECTOR DATABASES FOR RAGwhere the index sits, what it does on ingest and on query, and which one to pickINGEST (OFFLINE)Chunksfrom your chunking stepEmbedding modeltext → vectorQUERY (ONLINE)Queryuser questionEmbedsame model as ingestVector databaseindex: HNSW / IVFFlatvectors + ids + metadataapproximate nearest-neighbour searchfilter by metadata · optional hybrid (BM25 + vector)Top-k similaritynearest chunks by distanceRetrievedcontext → LLMThe main optionspgvectorChromaQdrantWeaviateMilvusPineconeSelf-hosted, open sourcepgvector · Chroma · Qdrant · Weaviate · Milvus — you run it, you tune itFully managedPinecone — plus cloud tiers of most OSS optionsPrototype on pgvector or Chroma · scale + filtering on Qdrant / Weaviate / Milvus · zero-ops on a managed service

What a vector database does in a RAG pipeline

In a retrieval-augmented generation pipeline, the vector database is the store that sits between your chunking step and the language model. It does two jobs. At ingest time, each chunk is run through an embedding model to produce a vector, and that vector is written to the database together with an id, the original text, and whatever metadata you attach (source file, section, tenant, date). At query time, the user's question is embedded with the same model, and the database returns the top-k stored vectors closest to it under a distance metric such as cosine, inner product, or L2. Those k chunks are the "context" that gets pasted into the prompt.

The part that makes a vector database more than a table with a float array column is the approximate nearest-neighbour (ANN) index. Exact search compares the query against every stored vector, which is fine for thousands of chunks and painful for millions. ANN indexes like HNSW (a graph-based index) and IVFFlat (a clustering-based index) trade a little recall for a large speed-up, and every option in this guide is built around one or more of them (pgvector README — indexing; Qdrant docs — indexing).

Two consequences follow. First, the database can only return what you indexed, so chunking quality caps retrieval quality before the database is ever involved. Second, the database's filtering and index behaviour directly shape what "top-k" means in practice — which is why the choice matters more than "it stores vectors" suggests.

Do you actually need a dedicated vector database?

This is the decision most teams get wrong in one of two directions: either they stand up a new distributed system for a corpus of ten thousand chunks, or they bolt vectors onto an existing database and only discover its limits under production filtering load. The honest answer is that a Postgres extension is enough for a lot of RAG apps, and a dedicated database earns its place under specific, nameable conditions.

The case for pgvector. pgvector adds a vector column type to Postgres, performs exact nearest-neighbour search by default (perfect recall), and lets you add an HNSW or IVFFlat index for approximate search when the table grows. It supports L2, inner product, and cosine distance, and — because it is just Postgres — you keep ACID transactions, joins against your existing tables, point-in-time recovery, and every client library you already use (pgvector README). If your chunks live next to the rows they came from, a single SELECT ... ORDER BY embedding <=> $1 LIMIT 10 with a WHERE clause on tenant or document id is a complete retrieval layer with no second system to operate.

A dedicated vector database earns its keep when:

  • • Scale. Your vector count is heading past what a single Postgres node's memory and HNSW build time handle comfortably, and you need horizontal sharding or a distributed deployment designed around vectors (Milvus's distributed mode, for example — Milvus docs).
  • • Filtering at scale. Most real queries combine "nearest to this vector" with "and only from this tenant / date range / document type". Dedicated databases design their ANN index around payload filtering so the filter doesn't collapse recall; Qdrant's filterable HNSW is the canonical example (Qdrant docs — filtering).
  • • Hybrid search. You want keyword (BM25) and vector scores fused in one query, which Weaviate ships natively (Weaviate docs — hybrid search). In Postgres you can combine tsvector full-text search with pgvector, but you are assembling the fusion yourself.
  • • Index tuning. You need to pick between HNSW, IVF, product quantisation, disk-based, or GPU indexes per collection and tune them independently of the rest of your database workload.
  • • Zero ops. Nobody on the team wants to run or tune a database at all, and a managed service is worth the vendor dependency.

If none of those apply yet, start with pgvector (or Chroma for a local prototype) and keep the retrieval interface behind a thin function, so swapping the store later is a contained change rather than a rewrite.

The main options compared

Below are the six options that come up in almost every "best vector database for RAG" conversation. The table sticks to documented characteristics — licence, hosting model, and what each is designed for — rather than benchmark rankings, which vary wildly with dataset, dimensionality, filter selectivity, and hardware.

OptionTypeHostingBest for
pgvectorPostgres extension (open source, PostgreSQL licence)Wherever your Postgres runs — self-hosted or any managed Postgres that ships the extensionTeams already on Postgres who want vectors next to their relational data, with SQL joins and ACID transactions, without a second system
ChromaOpen-source vector database (Apache 2.0), embedded / local-firstIn-process with a persistent client, self-hosted server, or Chroma CloudPrototyping and small-to-mid apps where a pip install and a local folder is the whole deployment story
QdrantOpen-source vector database (Apache 2.0), written in RustSelf-host (Docker / Kubernetes) or Qdrant CloudProduction workloads that lean hard on metadata filtering alongside vector search, with an option to move to managed later
WeaviateOpen-source vector database (BSD-3-Clause) with a managed cloudSelf-host or Weaviate CloudTeams that want built-in hybrid (BM25 + vector) search and pluggable vectoriser modules out of the box
MilvusOpen-source vector database (Apache 2.0), LF AI & Data projectMilvus Lite (pip), Standalone (Docker), Distributed (Kubernetes), or Zilliz CloudLarge-scale deployments that need a distributed architecture and a wide choice of index types, including disk- and GPU-based ones
PineconeFully managed, closed-source vector databaseManaged service only — no self-host optionTeams that want zero infrastructure to run and are comfortable with a vendor-hosted, proprietary store

A few details worth knowing beyond the table. Chroma can run entirely in-process (ephemeral or persisted to a directory) or as a client-server deployment, and its query API supports metadata filters via where and document-text filters via where_document (Chroma docs). Qdrant stores a JSON "payload" alongside each vector and supports sparse vectors and quantisation to shrink memory (Qdrant docs). Weaviate offers HNSW, flat, and dynamic index types plus vectoriser modules that call an embedding provider for you at import time (Weaviate docs). Milvus exposes the broadest index menu — FLAT, IVF variants, HNSW, DiskANN, and GPU indexes — and three deployment modes from a pip-installable Lite to a Kubernetes cluster (Milvus docs). Pinecone is serverless and managed only, with namespaces for tenant isolation and metadata filtering on query (Pinecone docs).

Get the weekly AI engineering brief

RAG, agents, evals, and the tools worth using — one practical email a week. Plus the free roadmap PDF.

By subscribing you agree to receive emails from AI Engineer Insights. Unsubscribe anytime. See our Privacy Policy.

Open source vs managed: the real trade-off

Five of the six options above are open-source vector databases or extensions, and four of those (Chroma, Qdrant, Weaviate, Milvus via Zilliz) also sell a managed cloud tier. That makes the choice less "open source vs managed" than "who runs it, and what do you give up either way".

Self-hosting an open source vector database gives you control over index parameters, data locality (the vectors never leave your VPC), and predictable infrastructure cost at steady load. The price is operations: you own upgrades, backups, replication, memory sizing for HNSW graphs, and re-indexing when you change embedding models. For a team that already runs Postgres or Kubernetes this is often marginal work; for a two-person team shipping a product it can be the largest single chunk of non-feature effort.

A managed service — Pinecone, or the cloud tier of an open-source option — removes that operational load entirely and usually scales without a re-architecture. What you give up is control and portability: index internals are the vendor's, cost scales with usage rather than with hardware, and migrating out means re-embedding or bulk-exporting your corpus. Choosing a managed tier of an open-source database keeps an exit hatch open, since the same API is available self-hosted; a closed-source service does not.

A reasonable rule: prototype on something you can run locally, and make the self-host vs managed call when you know your actual query volume, filter patterns, and who will be on-call.

A minimal vector database example for RAG (Chroma)

Here is the entire chunk → embed → store → query loop using the Chroma vector database, because it needs nothing but pip install chromadb and runs in-process. The same four calls — create a client, get a collection, add documents, query — map almost one-to-one onto every other option in this guide (Chroma docs — getting started).

import chromadb

# 1. Persistent client: vectors are stored on disk at ./rag_db and survive restarts
client = chromadb.PersistentClient(path="./rag_db")

# 2. A collection is a named index. Chroma embeds documents with its
#    default embedding function unless you pass embedding_function=...
collection = client.get_or_create_collection(name="docs")

# 3. Chunks from your chunking step (one string per chunk)
chunks = [
    "pgvector adds a vector type plus HNSW and IVFFlat indexes to Postgres.",
    "Chroma can run in-process with a persistent client, no server needed.",
    "Qdrant is an open-source vector database written in Rust.",
]

# 4. Add: each chunk is embedded, indexed, and stored with an id and metadata
collection.add(
    ids=[f"chunk-{i}" for i in range(len(chunks))],
    documents=chunks,
    metadatas=[{"source": "notes.md", "chunk": i} for i in range(len(chunks))],
)

# 5. Query: the question is embedded the same way, then the top-k nearest
#    chunks come back. 'where' applies an optional metadata filter.
results = collection.query(
    query_texts=["Which option runs inside Postgres?"],
    n_results=2,
    where={"source": "notes.md"},
)

for doc, dist in zip(results["documents"][0], results["distances"][0]):
    print(f"{dist:.3f}  {doc}")

Three things to notice. You never called an embedding model directly — Chroma applies a default embedding function to documents on both add and query, and you can swap in OpenAI, Sentence Transformers, or your own function by passing embedding_function= when you create the collection. Whatever you choose, ingest and query must use the same model or the distances are meaningless. Metadata is not optional in practice — the where filter is how you scope retrieval to a tenant, a document, or a date range, and it is the feature you will lean on most as the corpus grows. The distances come back with the documents, which is what you will log when you start measuring retrieval quality in the next section.

Swapping this for pgvector means a CREATE EXTENSION vector, a table with a vector(n) column, an INSERT per chunk, and an ORDER BY embedding <=> query_vector LIMIT k query — with the embedding call done in your own code (pgvector README). For Qdrant, Weaviate, Milvus, and Pinecone the shape is the same: a client, a collection/index, an upsert of vectors plus payload, and a search with a filter.

How to choose: a decision guide by stage

  • • Prototype / local development. Chroma with a persistent client, or pgvector if you already have Postgres running. Optimise for iteration speed — you will change chunking and embedding models several times, and re-indexing a local store is free.
  • • Already on Postgres, up to millions of vectors. pgvector with an HNSW index. You get joins against your application data, one backup story, and no new system. Revisit when index build time, memory, or filtered-query recall becomes a measurable problem rather than a hypothetical one.
  • • Production with heavy metadata filtering or hybrid search. Qdrant (filterable HNSW, payloads, sparse vectors), Weaviate (native hybrid search, vectoriser modules), or Milvus (distributed deployment, wide index choice). All three are open source, so you can self-host now and move to their managed cloud later without changing the client code.
  • • Zero-ops, willing to accept a vendor. Pinecone, or the managed tier of one of the open-source databases. Pick the managed tier of an open-source option if you want the ability to self-host later; pick Pinecone if you never intend to.
  • • Very large scale (hundreds of millions of vectors and up). Milvus Distributed or a managed service designed for that scale. This is the only tier where a dedicated distributed architecture is a requirement rather than a preference.

Whatever you pick, put the retrieval call behind a single function that takes a query string and filter and returns chunks with scores. Every option here fits that interface, and it is the difference between a one-afternoon migration and a week of untangling.

How to know your vector DB retrieval is actually good

The database returns top-k vectors by distance — that is a mechanical guarantee, not a quality one. Whether those k chunks are the right chunks depends on the embedding model, the chunking, the index's recall at your settings, and how well filters interact with the index. None of that is visible from the query response, so treat the store the way you'd treat any other pipeline component: hold everything else constant and measure retrieval.

The two retrieval-side metrics to track are context precision (of the chunks returned, how many were relevant) and context recall (of the chunks that were relevant, how many were returned), computed against a labelled eval set of questions and their expected source chunks. Frameworks such as RAGAS implement both (RAGAS docs), and our guide to RAG evaluation metrics walks through how they're calculated and what thresholds teams commonly use.

Two database-specific checks are worth adding. First, compare ANN recall against exact search on a sample: pgvector lets you drop the index (or raise hnsw.ef_search) and re-run the same query to see how much recall the index is costing you (pgvector README — query options); Qdrant and the others expose equivalent search-time parameters. Second, re-run your eval set with your most selective production filter applied — a store that scores well unfiltered and badly filtered is telling you something important about how its index handles that filter. A vector database is only as good as the retrieval quality you can measure through it.

Frequently Asked Questions

What is a vector database in RAG?

A vector database is the store that holds the embeddings of your document chunks and answers nearest-neighbour queries against them. In a RAG pipeline it sits between chunking and generation: at ingest time each chunk is embedded and written to the database with an id and metadata, and at query time the user's question is embedded with the same model and the database returns the top-k most similar chunks, which become the context the language model answers from.

Do I need a vector database for RAG?

Not always. For small-to-mid scale — up to a few million vectors — pgvector inside Postgres is often enough, especially if you already run Postgres and want vectors next to your relational data. A dedicated vector database earns its place when you need to scale beyond what a single Postgres node handles comfortably, when you rely heavily on metadata filtering combined with vector search, or when you want hybrid (keyword plus vector) search and finer control over the ANN index.

What is the best open-source vector database?

There is no single winner — it depends on where you are. Chroma is the easiest to start with because it runs in-process from a pip install and persists to a local folder, which makes it ideal for local development and prototyping. Qdrant, Weaviate, and Milvus are all open-source options built for self-hosted production: Qdrant is known for filtering alongside vector search, Weaviate for built-in hybrid search and vectoriser modules, and Milvus for a distributed architecture and a wide choice of index types. Pick based on your deployment model and the features you will actually use, not on a leaderboard.

Is pgvector good enough for RAG?

Yes, for many production applications. pgvector adds a vector column type, exact nearest-neighbour search by default, and HNSW and IVFFlat indexes for approximate search, and it inherits Postgres transactions, joins, and backups. It is a strong choice up to millions of vectors, particularly if you are already on Postgres. The caveats are that ANN index build time and memory (especially for HNSW) grow with the table, and that scaling beyond one node is a Postgres scaling problem rather than a vector-database feature, so very large or very filter-heavy workloads may be better served by a dedicated store.

Chroma vs Pinecone: which should I use for RAG?

They solve different problems. Chroma is open source and local-first: it runs in-process with a persistent client, so it is great for development, prototyping, and small self-hosted apps, and it also offers a server mode and a managed cloud. Pinecone is a fully managed, closed-source service: you never run infrastructure, it scales without you operating anything, and in return you accept a proprietary store and a vendor dependency. Use Chroma to build and iterate; move to Pinecone (or a managed tier of an open-source database) when you want someone else to run production.

How do I know if my vector database is returning good results?

Measure retrieval directly rather than judging the final answers by eye. Build a small eval set of questions with the chunks that should be retrieved, then compute context precision (how much of what was retrieved is relevant) and context recall (how much of what was relevant was retrieved) for each configuration — index type, top-k, filters, embedding model. Our guide to RAG evaluation metrics covers how those scores are computed. The database choice only matters to the extent that retrieval quality you can measure improves.

Want 1:1 help? Book a session

5.0 · 18 reviews

Career guidance, resume & interview prep, or tech consulting — 1:1 with a Lead AI Engineer. Start with a free 30-min quick chat.

References

Found this useful? Share it.

Share:

Related Articles