GenAI System Design
3. Embeddings and Vector Search

Embeddings for System Designers

What embeddings are, how dimensions and precision drive storage and cost, which similarity metric to use, and the operational traps of embedding pipelines.

Lesson 1 of 7 9 min

What an embedding is

An embedding model turns a piece of content into a list of numbers, a vector, typically 384 to 3,072 long. It is trained so that content with similar meaning lands close together in that space. "How do I reset my password?" and "I forgot my login credentials" get nearby vectors despite sharing no keywords.

Embeddings power semantic search, RAG retrieval, recommendations, deduplication, clustering and classification. For a system designer, an embedding is just a fixed-size, dense feature vector, and the design questions are about producing, storing and searching billions of them.

Similarity metrics

Cosine similarity, dot product and distance

Drag the arrow tips. Real embeddings have hundreds of dimensions, but the geometry is the same.

querydoc
Cosine similarity
0.780
angle 39°, ignores length
Dot product
1.270
angle and length
Euclidean distance
0.849
smaller = closer
Lengths
1.30 / 1.25
dashed circle = unit length

Drag both tips onto the dashed circle: when vectors are normalised to length 1, cosine, dot product and distance all rank neighbours the same way. That is why most embedding APIs return unit vectors and databases use the cheaper dot product.

MetricFormula (intuition)Notes
Cosine similarityangle between vectorsIgnores length. The default for text embeddings
Dot productlength × length × cos(angle)Equals cosine when vectors are unit length, and is cheaper
Euclidean (L2) distancestraight-line distanceSame ranking as cosine for unit vectors

Most embedding APIs return normalised (unit-length) vectors, so all three agree and databases use the cheapest one, the dot product. Use the metric the embedding model was trained with, which its documentation will state.

Dimensions, precision and cost

Bytes per vector = dimensions × bytes per dimension

Dimensionsfloat32float16int8binary (1 bit)
3841.5 KB768 B384 B48 B
7683 KB1.5 KB768 B96 B
1,5366 KB3 KB1.5 KB192 B
3,07212 KB6 KB3 KB384 B

At 100M vectors, 1,536-dim float32 is 614 GB, while int8 is 154 GB and binary is 19 GB. Memory is the main cost driver of a vector database, so these choices matter more than which vendor you pick.

Two modern features help:

  • Matryoshka embeddings. Many recent models are trained so that the first N dimensions are a usable embedding on their own. You can store 256 dims for fast first-pass search and keep the full 1,536 for rescoring.
  • Quantized embeddings. int8 or binary quantization of vectors, often with a rescoring step using full precision, keeps most of the recall at 4–32× less memory.

Choosing an embedding model

  • Quality on your data. Public leaderboards like MTEB are a starting point. Test on a sample of your own queries and documents.
  • Domain and language. Code, legal, medical and multilingual content often benefit from specialised models.
  • Maximum input length. Many models accept 512–8,192 tokens. Longer chunks are truncated silently.
  • Hosted vs self-hosted. API embedding costs are small per token, around $0.02–0.15 per million, but add up over billions of chunks. Self-hosted small models (100M–600M parameters) embed thousands of chunks per second on one GPU.
  • Asymmetric use. Many models expect different prefixes or modes for queries and documents. Using the wrong one quietly hurts recall.

The embedding pipeline

Source change
doc created, edited, deleted
Parse & chunk
text + metadata
Embed
batched, GPU or API
Upsert
vector + ACL + metadata
Index
HNSW / IVF build
  • Batch embedding calls. Embedding models are small and fast when batched, and slow per request.
  • Store the model name and version with every vector.
  • Handle deletes and edits. Stale vectors pointing at deleted documents cause wrong answers and compliance problems.

Key takeaways

  • An embedding model maps text (or images) to a fixed-length vector where similar meaning means nearby vectors.
  • Storage per vector = dimensions × bytes per dimension. 1,536 float32 dims is 6 KB, so 100M vectors is 600 GB raw.
  • With normalised vectors, cosine similarity, dot product and Euclidean distance give the same ranking. Use dot product.
  • Vectors from different embedding models are not comparable. Changing models means re-embedding everything.

Go deeper

Finished reading? Mark it done to track your progress.