Embeddings for System Designers
What embeddings are, how dimensions and precision drive storage and cost, which similarity metric to use, and the operational traps of embedding pipelines.
What an embedding is
An embedding model turns a piece of content into a list of numbers, a vector, typically 384 to 3,072 long. It is trained so that content with similar meaning lands close together in that space. "How do I reset my password?" and "I forgot my login credentials" get nearby vectors despite sharing no keywords.
Embeddings power semantic search, RAG retrieval, recommendations, deduplication, clustering and classification. For a system designer, an embedding is just a fixed-size, dense feature vector, and the design questions are about producing, storing and searching billions of them.
Similarity metrics
Cosine similarity, dot product and distance
Drag the arrow tips. Real embeddings have hundreds of dimensions, but the geometry is the same.
Drag both tips onto the dashed circle: when vectors are normalised to length 1, cosine, dot product and distance all rank neighbours the same way. That is why most embedding APIs return unit vectors and databases use the cheaper dot product.
| Metric | Formula (intuition) | Notes |
|---|---|---|
| Cosine similarity | angle between vectors | Ignores length. The default for text embeddings |
| Dot product | length × length × cos(angle) | Equals cosine when vectors are unit length, and is cheaper |
| Euclidean (L2) distance | straight-line distance | Same ranking as cosine for unit vectors |
Most embedding APIs return normalised (unit-length) vectors, so all three agree and databases use the cheapest one, the dot product. Use the metric the embedding model was trained with, which its documentation will state.
Dimensions, precision and cost
Bytes per vector = dimensions × bytes per dimension
| Dimensions | float32 | float16 | int8 | binary (1 bit) |
|---|---|---|---|---|
| 384 | 1.5 KB | 768 B | 384 B | 48 B |
| 768 | 3 KB | 1.5 KB | 768 B | 96 B |
| 1,536 | 6 KB | 3 KB | 1.5 KB | 192 B |
| 3,072 | 12 KB | 6 KB | 3 KB | 384 B |
At 100M vectors, 1,536-dim float32 is 614 GB, while int8 is 154 GB and binary is 19 GB. Memory is the main cost driver of a vector database, so these choices matter more than which vendor you pick.
Two modern features help:
- Matryoshka embeddings. Many recent models are trained so that the first N dimensions are a usable embedding on their own. You can store 256 dims for fast first-pass search and keep the full 1,536 for rescoring.
- Quantized embeddings. int8 or binary quantization of vectors, often with a rescoring step using full precision, keeps most of the recall at 4–32× less memory.
Choosing an embedding model
- Quality on your data. Public leaderboards like MTEB are a starting point. Test on a sample of your own queries and documents.
- Domain and language. Code, legal, medical and multilingual content often benefit from specialised models.
- Maximum input length. Many models accept 512–8,192 tokens. Longer chunks are truncated silently.
- Hosted vs self-hosted. API embedding costs are small per token, around $0.02–0.15 per million, but add up over billions of chunks. Self-hosted small models (100M–600M parameters) embed thousands of chunks per second on one GPU.
- Asymmetric use. Many models expect different prefixes or modes for queries and documents. Using the wrong one quietly hurts recall.
The embedding pipeline
- Batch embedding calls. Embedding models are small and fast when batched, and slow per request.
- Store the model name and version with every vector.
- Handle deletes and edits. Stale vectors pointing at deleted documents cause wrong answers and compliance problems.
Key takeaways
- An embedding model maps text (or images) to a fixed-length vector where similar meaning means nearby vectors.
- Storage per vector = dimensions × bytes per dimension. 1,536 float32 dims is 6 KB, so 100M vectors is 600 GB raw.
- With normalised vectors, cosine similarity, dot product and Euclidean distance give the same ranking. Use dot product.
- Vectors from different embedding models are not comparable. Changing models means re-embedding everything.