GenAI System Design
10. GenAI System Design Problems

Designing Enterprise Document Q&A

A full walkthrough for an internal assistant that answers questions over Confluence, Drive, Slack and Jira for 50K employees, with permissions, freshness, citations and retrieval quality at its core.

Lesson 2 of 10 15 min

1. Requirements

Functional

  • Employees ask natural-language questions and get answers with citations to source documents
  • Sources: Confluence, Google Drive, Slack (public channels plus the user's private ones), Jira
  • Follow-up questions in a conversation
  • Answers must only use content the asking user can access in the source system

Non-functional

  • 50K employees, around 20M documents
  • Changes searchable within 15 minutes. Deletions and permission changes within 15 minutes
  • First token under 2 s at p95
  • Auditable: who asked what, and which documents were used

2. Operating point

Quality and trust first. Wrong or leaked answers destroy adoption. Latency of 1–2 s to first token is fine for an internal tool, and cost is moderate at this scale.

3. Estimates

QuantityValue
Questions per day50K × 5 = 250K
Peak QPS (work hours, 3×)≈ 9 QPS
Chunks20M docs × ≈ 5 chunks = 100M chunks
Vector index (1,024-dim int8, HNSW M=16)100M × (1 KB + ≈ 140 B) ≈ 115 GB per replica
Daily changes (≈ 2% of docs)400K docs, so ≈ 2M chunks re-embedded per day
Tokens per question≈ 6K in (instructions + 8 chunks + history), 400 out
LLM cost ($2 / $10 per M)250K × ($0.012 + $0.004) ≈ $4K a day, ≈ $120K a month before caching

At 9 QPS, a hosted API is clearly the right call for generation. Embeddings can use an API or one small self-hosted GPU.

4. Architecture

Ingestion (offline)

Connectors
webhooks + incremental crawl
Change queue
per-doc events
Parse + chunk
structure-aware
Embed
batched
Hybrid index
vectors + BM25 + ACLs

A separate ACL sync service pulls permissions (document shares, space permissions, channel membership) and group memberships from each source and the identity provider, and updates chunk metadata.

Query (online)

User (SSO)
Identity + groups
cached expansion
Query rewrite
standalone + filters
Filtered hybrid search
top 100, ACL filter
Rerank → 8
LLM + citations
verify cited IDs

5. Deep dives

Deep dive A: permissions

  • Model: each chunk stores source, allowed_users, allowed_groups and visibility. Large groups are stored as group IDs, not expanded user lists.
  • Query-time identity: resolve the user's groups (including nested groups) from the identity provider, cached for about 5 minutes, and add the filter visibility = public OR groups ∩ user_groups ≠ ∅ OR user ∈ allowed_users.
  • Late-binding check for the final 8 chunks against the source system's permission API, where available. This catches ACL-sync lag for the few documents actually shown.
  • Slack private channels and DMs: index them only for their members, or exclude DMs entirely by policy.
  • Everything else is scoped too: caches keyed by permission scope, and traces restricted to authorised reviewers.

Deep dive B: freshness and deletes

  • Webhooks give most updates within seconds. Sources without webhooks are polled with "modified since" every 5 minutes.
  • Nightly reconciliation: list IDs per source, diff them against the index, and fix missing, stale and deleted documents.
  • Versioned chunk replacement: a new version atomically replaces all of a document's chunks. Deletes tombstone immediately and compact later.
  • Permission changes flow as high-priority events. They are more urgent than content edits.

Deep dive C: retrieval quality

  • Parsing tuned per source: Confluence macros, Drive tables, Slack threads grouped into conversation chunks, Jira issue fields.
  • Hybrid search (project codes, error messages and names need BM25), plus a reranker.
  • Recency boost for fast-changing sources (Slack, Jira), and authority boost for official spaces.
  • Eval set: 500 real questions collected with each department, labelled with the correct source documents. Track recall@20 for retrieval and answer correctness with an LLM judge on every change.

6. Evaluation and safety

  • Faithfulness checks: verify citations exist in the provided context, and flag unsupported claims.
  • "Don't know" behaviour: when rerank scores are low, say so and suggest where to look. Log these as content gaps.
  • Prompt injection: documents are untrusted input. The assistant has no write tools, which limits the damage to wrong answers within the user's own access.
  • Audit log: user, question, retrieved document IDs and answer, with retention per policy.

7. Bottlenecks and follow-ups

Likely questionAnswer sketch
"A user got an answer from a doc they can't see?"ACL-sync lag. Add a late-binding check, prioritise permission events, alert on reconciliation diffs
"Answers are stale"Measure freshness lag per connector, fix polling gaps, add a recency boost
"Scale to 500K employees?"The index grows about 10× (shard it), QPS to about 90. Still API-friendly. Consider a per-tenant index split by business unit
"Multi-hop questions?"Route to an agentic retrieval path with a step budget

Key takeaways

  • At this scale the hard problems are permissions, freshness and retrieval quality, not GPU capacity.
  • Sync ACLs from every source, attach them to every chunk, and filter retrieval by the caller's identity and groups.
  • Event-driven connectors plus a reconciliation job keep the index within minutes of the sources, including deletes.
  • Use hybrid retrieval, a reranker and citations, measured on a labelled question set built with each department.

Go deeper