Ingestion Pipelines and Chunking
How to turn messy source documents into retrievable chunks, including parsing, chunking strategies and sizes, metadata, and the pipeline architecture that keeps it running at scale.
Parsing: the unglamorous first step
Real corpora are PDFs, slide decks, HTML pages with navigation noise, spreadsheets, scanned images and wiki markup. Before chunking, you need clean text with structure preserved:
- Headings and hierarchy. They become metadata and chunk boundaries.
- Tables. Convert them to Markdown or row-wise sentences ("Plan: Pro, Price: $20, Seats: 5"). Flattened tables are a top source of wrong answers.
- Boilerplate removal. Headers, footers, navigation and cookie banners repeated on every page pollute retrieval.
- OCR or vision models for scanned documents and images, and layout-aware parsers for complex PDFs.
Spend time here. Inspect a random sample of parsed output before you build anything else.
Chunking strategies
Chunking playground
Split a refund policy three ways. The question to answer later: "When does the refund window start for enterprise customers?"
- Chunk 1≈ 48 tokens
# Refund policy ## Eligibility Customers on annual plans can request a full refund within 30 days of purchase. Monthly plans are not refundable, but you can cancel at any time and keep access
- Chunk 2≈ 48 tokens
until the end of the billing period. ## Enterprise contracts Enterprise customers follow the terms in their signed order form. If the order form is silent on refunds, the standard 30-day win
- Chunk 3≈ 48 tokens
dow applies from the contract start date, not the invoice date. ## How to request Open a ticket from the billing page and choose "Refund request". Refunds are issued to the original payment m
- Chunk 4≈ 48 tokens
ethod within 5–10 business days. Bank transfers can take up to 15 business days. ## Exceptions We do not refund add-on usage such as extra API calls or storage overages, even within the 30-da
- Chunk 5≈ 3 tokens
y window.
Fixed-size chunks cut sentences in half and separate facts from the heading that gives them meaning. Structure-aware chunking keeps sentences whole and repeats the section title in each chunk, so "30 days from the contract start date" stays attached to "Enterprise contracts".
| Strategy | How | Good for |
|---|---|---|
| Fixed size (+ overlap) | Every N tokens, with M tokens of overlap | Quick baseline. Uniform text |
| Recursive / structure-aware | Split by headings, then paragraphs, then sentences, merging up to a size limit | The practical default for documents |
| Section-based | One chunk per heading section | Well-structured docs with short sections |
| Semantic | Split where embedding similarity between adjacent sentences drops | Long unstructured text |
| Special formats | Code by function or class, chats by thread, tables by row group | Non-prose content |
Chunk size
There is a trade-off:
- Small chunks (100–300 tokens): precise matches, but they lose surrounding context, and you need more of them in the prompt.
- Large chunks (800–1,500 tokens): carry context, but embeddings get blurry because one vector averages many topics, and each chunk spends more of the prompt.
Pick by measurement: build a small eval set of questions with known source passages, then compare recall@k across chunking settings.
Making chunks self-contained
A chunk saying "This applies from the contract start date" is useless without knowing what "this" is. Techniques:
- Prepend context: document title plus section path ("Refund policy > Enterprise contracts") on every chunk.
- Contextual enrichment: use an LLM at ingestion time to write a one-sentence summary of where the chunk sits in the document, and prepend it before embedding. This adds ingestion cost (one small LLM call per chunk), but measurably improves retrieval. Prompt caching of the full document makes it cheaper.
- Parent-child (small-to-big) retrieval: index small chunks for precise matching, but return their larger parent section to the LLM.
Metadata
Store with every chunk:
doc_id,chunk_index,versionor hash: for updates and deletessource,url,title,section: for citationscreated_at,updated_at: for recency filters and boosts- Access control: tenant, groups, visibility (permissions lesson)
embedding_model: for migrations
Pipeline architecture
- Idempotent by
(doc_id, version). Reprocessing the same version is a no-op, and a new version replaces all of that document's old chunks. - Skip unchanged content by hashing each chunk. Only re-embed chunks whose text changed.
- Backpressure and retries. Embedding APIs rate-limit, so queue and retry with backoff.
- Dead-letter queue for documents that fail to parse, with alerting. Silent drops become missing-knowledge bugs.
- Backfill mode for re-embedding everything, with its own throughput limits, so it doesn't starve live updates.
Key takeaways
- Garbage in, garbage out. Parsing quality (tables, headings, PDFs) often matters more than the embedding model.
- Chunk along document structure, keep sentences whole, and attach the section title and document context to each chunk.
- Typical chunk sizes are 200–800 tokens. Smaller improves precision, larger keeps context. Measure on your own eval set.
- Run ingestion as an event-driven, idempotent pipeline keyed by document ID and version.