GenAI System Design
3. Embeddings and Vector Search

Choosing a Vector Store

pgvector, search engines with vector support, dedicated vector databases and in-process libraries, compared on scale, filtering, operations and cost.

Lesson 5 of 7 9 min

The four options

In-process library
FAISS, hnswlib, USearch
Postgres + pgvector
vectors next to your rows
Search engine
Elasticsearch, OpenSearch, Vespa
Vector database
Pinecone, Qdrant, Milvus, Weaviate, turbopuffer
Rough order of operational weight, from embedded in your process to a dedicated distributed service.

In-process libraries

FAISS, hnswlib and USearch run inside your application process. They are extremely fast and free, but you own persistence, replication, updates and sharding. Use them for batch jobs, offline evaluation, read-only indexes shipped with a service, or as the engine inside something you build.

Postgres + pgvector

pgvector adds a vector column type and HNSW or IVFFlat indexes to PostgreSQL.

  • Pros: vectors live next to your relational data, with transactions, joins, backups and access control you already have. Filtering is plain SQL. There is no new system to operate.
  • Cons: index size is bounded by one Postgres node's memory unless you shard. Heavy vector search competes with OLTP traffic, and very large index builds are slow.
  • Sweet spot: up to roughly tens of millions of vectors, especially multi-tenant SaaS where every query filters by tenant.

Search engines with vector support

Elasticsearch, OpenSearch, Vespa and Solr now have ANN indexes beside their inverted indexes.

  • Pros: hybrid search (BM25 + vectors) in one query, mature filtering, facets and aggregations, and proven distributed operations.
  • Cons: heavier to run. Vector performance per GB of RAM is often behind specialised engines.
  • Sweet spot: you already run search, or keyword relevance matters as much as semantic relevance (e-commerce, docs search).

Dedicated vector databases

Pinecone, Qdrant, Milvus, Weaviate, turbopuffer, Chroma and others.

  • Pros: built for billions of vectors. They offer quantization, disk-based indexes, in-index filtering, namespaces for multi-tenancy, and managed serverless options that separate storage (object storage) from compute.
  • Cons: another system to operate or pay for, and data must be synchronised from your source of truth.
  • Sweet spot: hundreds of millions of vectors or more, high QPS, or a team that wants a managed, vector-first service.

What actually differentiates them

Question to askWhy it matters
How is filtering done: pre, post, or in-index with a selectivity-aware planner?Wrong strategy means empty results or brute-force slowness
Native hybrid search and rank fusion?Most production RAG needs keyword plus vector search
Multi-tenancy: namespaces, partitions, per-tenant indexes?Isolation, noisy neighbours, per-tenant deletion (GDPR)
Updates and deletes: real-time? When is compaction needed?Freshness and index degradation
Memory model: all in RAM, quantized, disk-based, object-storage tiered?The dominant cost at scale
Consistency: read-your-writes after an upsert?Users expect a just-uploaded document to be searchable
Scaling: sharding, replication, rebalancing?Operability past one node

A decision rule for interviews

  1. Under ~10M vectors and you run Postgres: pgvector. Say so confidently, because simplicity is a feature.
  2. You need strong keyword relevance, or you already run a search cluster: use its vector support for hybrid search in one system.
  3. Hundreds of millions to billions of vectors, heavy filtering, high QPS: a dedicated vector database, managed if the team is small.
  4. Read-only or batch: an in-process library, built offline and shipped as an artifact.

Key takeaways

  • Start from what you already run. pgvector or your existing search engine handles millions of vectors with no new infrastructure.
  • Dedicated vector databases earn their place at hundreds of millions of vectors, with heavy filtering or high QPS.
  • Evaluate filtering, hybrid search, multi-tenancy, update and delete behaviour, and cost per GB of index, not just ANN benchmarks.
  • The vector store is rarely the bottleneck of a RAG system. Pick the simplest option that meets scale.

Go deeper

Finished reading? Mark it done to track your progress.