Freshness, Deletes and Re-indexing
Keeping a RAG index in sync with its sources, covering change detection, update and delete semantics, freshness SLOs, versioned documents, and safe re-indexing migrations.
Why freshness is a design problem
The model answers from whatever is in the index. If a policy changed yesterday and the index still has last month's version, the assistant confidently gives the old answer, with a citation. Stale RAG is worse than no RAG because it looks authoritative.
Agree on a freshness SLO early: "Changes are searchable within 5 minutes" is a very different system from "within 24 hours".
Detecting changes
| Method | Freshness | Cost | Notes |
|---|---|---|---|
| Webhooks / events from source apps | Seconds | Low | Best when sources support it (wikis, ticketing, storage buckets) |
| Change data capture (CDC) from databases | Seconds | Low | Debezium-style streams from the source DB's log |
| Polling with "modified since" | Minutes | Medium | Works with most APIs |
| Full re-crawl | Hours to days | High | Simple, and catches everything, including deletes |
Most production systems combine events for speed with a periodic reconciliation: list source IDs, compare them with index IDs, and fix drift. Events get lost, and reconciliation catches what they miss.
Update semantics
When a document changes:
- Re-parse and re-chunk the new version.
- Hash each chunk and re-embed only chunks whose content changed. This saves embedding cost for small edits to big documents.
- Replace the document's chunk set atomically: write new chunks tagged with the new version, then delete chunks from older versions.
Key chunks by (doc_id, chunk_index), or better by content hash, so you never end up with duplicate or orphaned chunks.
Deletes
Deletes matter for correctness and for compliance. A deleted document, or a user's data after a GDPR request, must stop appearing in answers.
- Propagate deletes through the same pipeline as updates, with high priority.
- Understand your vector store's delete behaviour. Many use tombstones plus background compaction, so data is hidden immediately but physically removed later. Know the compaction schedule when you promise deletion times.
- Remember the other copies: caches (semantic caches of old answers), logs and traces containing retrieved text, and backups.
Versioned and time-sensitive content
- Multiple valid versions: product docs for v3 and v4, policies per region. Store version and region as metadata and filter on them, instead of keeping only the latest.
- Effective dates: a policy that starts next month should not answer today's questions. Store
effective_fromandeffective_toand filter on them. - Recency boosting: for news, tickets or incident reports, boost newer chunks in ranking.
Re-indexing migrations
You will re-process the whole corpus when you:
- change embedding model or dimensions
- change chunking strategy
- change index type or parameters, or vector store
Treat it like a database migration:
Monitoring freshness
- Lag: time from a source change to searchable. Alert if it exceeds the SLO.
- Coverage: document count in the source compared with the index, by source.
- Failure queues: documents failing parse or embed, and their age.
- Staleness spot checks: sample recent source edits and verify that the index has the new version.
Key takeaways
- The index is a derived copy of your sources. Every create, update and delete must flow through to it, or answers go stale or leak deleted data.
- Prefer event-driven change capture (webhooks, CDC) over periodic full crawls, with a reconciliation job as a safety net.
- Replace all chunks of a document atomically on update, keyed by document ID and version.
- Re-embedding and re-chunking are routine migrations. Build them as blue/green index swaps.