GenAI System Design
4. RAG Architecture

Freshness, Deletes and Re-indexing

Keeping a RAG index in sync with its sources, covering change detection, update and delete semantics, freshness SLOs, versioned documents, and safe re-indexing migrations.

Lesson 4 of 7 9 min

Why freshness is a design problem

The model answers from whatever is in the index. If a policy changed yesterday and the index still has last month's version, the assistant confidently gives the old answer, with a citation. Stale RAG is worse than no RAG because it looks authoritative.

Agree on a freshness SLO early: "Changes are searchable within 5 minutes" is a very different system from "within 24 hours".

Detecting changes

MethodFreshnessCostNotes
Webhooks / events from source appsSecondsLowBest when sources support it (wikis, ticketing, storage buckets)
Change data capture (CDC) from databasesSecondsLowDebezium-style streams from the source DB's log
Polling with "modified since"MinutesMediumWorks with most APIs
Full re-crawlHours to daysHighSimple, and catches everything, including deletes

Most production systems combine events for speed with a periodic reconciliation: list source IDs, compare them with index IDs, and fix drift. Events get lost, and reconciliation catches what they miss.

Update semantics

When a document changes:

  1. Re-parse and re-chunk the new version.
  2. Hash each chunk and re-embed only chunks whose content changed. This saves embedding cost for small edits to big documents.
  3. Replace the document's chunk set atomically: write new chunks tagged with the new version, then delete chunks from older versions.

Key chunks by (doc_id, chunk_index), or better by content hash, so you never end up with duplicate or orphaned chunks.

Deletes

Deletes matter for correctness and for compliance. A deleted document, or a user's data after a GDPR request, must stop appearing in answers.

  • Propagate deletes through the same pipeline as updates, with high priority.
  • Understand your vector store's delete behaviour. Many use tombstones plus background compaction, so data is hidden immediately but physically removed later. Know the compaction schedule when you promise deletion times.
  • Remember the other copies: caches (semantic caches of old answers), logs and traces containing retrieved text, and backups.

Versioned and time-sensitive content

  • Multiple valid versions: product docs for v3 and v4, policies per region. Store version and region as metadata and filter on them, instead of keeping only the latest.
  • Effective dates: a policy that starts next month should not answer today's questions. Store effective_from and effective_to and filter on them.
  • Recency boosting: for news, tickets or incident reports, boost newer chunks in ranking.

Re-indexing migrations

You will re-process the whole corpus when you:

  • change embedding model or dimensions
  • change chunking strategy
  • change index type or parameters, or vector store

Treat it like a database migration:

Build index v2
in parallel
Backfill
batch re-embed, rate-limited
Dual-write
live changes → v1 and v2
Evaluate
recall + answer quality on eval set
Switch alias
keep v1 for rollback

Monitoring freshness

  • Lag: time from a source change to searchable. Alert if it exceeds the SLO.
  • Coverage: document count in the source compared with the index, by source.
  • Failure queues: documents failing parse or embed, and their age.
  • Staleness spot checks: sample recent source edits and verify that the index has the new version.

Key takeaways

  • The index is a derived copy of your sources. Every create, update and delete must flow through to it, or answers go stale or leak deleted data.
  • Prefer event-driven change capture (webhooks, CDC) over periodic full crawls, with a reconciliation job as a safety net.
  • Replace all chunks of a document atomically on update, keyed by document ID and version.
  • Re-embedding and re-chunking are routine migrations. Build them as blue/green index swaps.

Go deeper

Finished reading? Mark it done to track your progress.