GenAI System Design
10. GenAI System Design Problems

Designing an AI Coding Assistant

A full walkthrough for a Copilot-style assistant with inline completions and chat for 1M developers, covering the latency budget, prefill-heavy serving, context selection, repo retrieval, privacy, and acceptance-based evaluation.

Lesson 3 of 10 15 min

1. Requirements

Functional

  • Inline completions as the developer types (fill-in-the-middle)
  • Chat about the codebase: explain, refactor, write tests
  • Agent mode: multi-file edits, run tests, iterate
  • IDE plugins for major editors

Non-functional

  • 1M active developers
  • Inline suggestions within about 300 ms end to end, or they are useless
  • Code privacy: no retention or training on customer code by default, and enterprise controls
  • Cost per developer well under the subscription price

2. Operating point

Two operating points in one product:

  • Inline: latency above everything, then cost. Quality needs to be good enough that developers accept them.
  • Chat and agent: quality first, with latency of seconds (or minutes for agents) acceptable.

3. Estimates (inline completions)

QuantityValue
Completion requests per dev per day (after debounce)≈ 1,000
Requests per day1B
Average / peak QPS (3×)≈ 11.6K / ≈ 35K
Input / output tokens≈ 2K context / ≈ 30 tokens
Peak input tokens/s35K × 2K ≈ 70M tokens/s, prefill-dominated
Peak output tokens/s35K × 30 ≈ 1M tokens/s

This is the opposite of chat: prefill dominates. Each keystroke's request shares almost all of its prefix with the previous one (same file, same cursor region), so prefix caching with session affinity turns 2K tokens of prefill into about 100 new tokens per request. Without it, the fleet would need several times more GPUs.

At about $0.25 per million tokens for a small model via an API, uncached input alone would be 1B × 2K = 2T tokens a day, or about $500K a day. That is why completion models are small (1–15B parameters), self-hosted, and heavily cached.

4. Architecture

Inline path

IDE plugin
context gather, debounce, cancel
Edge proxy
auth, nearest region
Completion service
prompt assembly, filters
Inference pool
small code model, FIM
Post-filter
secrets, duplicates, syntax

Chat and agent path

IDE chat panel
Gateway
frontier model routing
Repo retrieval
symbols + embeddings + grep
Agent loop
edit, run, observe
Sandbox
tests, builds

5. Deep dives

Deep dive A: the inline latency budget

StepBudget
Debounce after the last keystroke≈ 50–100 ms (not counted as model latency)
Network to the nearest region20–50 ms
Prompt assembly< 10 ms
Model TTFT (cached prefix)50–120 ms
Generate ≈ 30 tokens at 200+ tokens/s≈ 100–150 ms
Total after debounce≈ 200–300 ms

Techniques:

  • Cancellation: a new keystroke cancels the in-flight request, freeing GPU work immediately.
  • Streaming partial suggestions, with the plugin showing the first line early.
  • Speculative decoding with n-gram lookup, because code repeats nearby text often.
  • Regional inference pools close to developers.
  • Client-side caching of recent suggestions, for backspace and retype patterns.

Deep dive B: context selection

A 2K-token budget is tiny compared with a repository. Rank candidate context:

  1. Code before and after the cursor (fill-in-the-middle prefix and suffix)
  2. Imports, and the signatures of referenced symbols (from the language server or ctags)
  3. Snippets from recently viewed and edited files, which strongly predict relevance
  4. Similar code from the repo (embedding search), for chat or larger budgets

Assemble the static parts (file header, imports) first to maximise prefix-cache reuse.

Deep dive C: privacy and safety

  • No retention of prompts and suggestions by default. Only aggregated telemetry (accept or reject, latency).
  • Secret filtering: block suggestions containing API keys, and avoid sending obvious secret files as context.
  • Licence and duplicate filters for suggestions matching public code verbatim, where required.
  • Enterprise controls: exclude repositories or paths, region pinning, a self-hosted option.
  • Agent sandbox: isolated containers or microVMs, network allow-list, no production credentials, and user approval before pushing or running destructive commands.

6. Evaluation

  • Online: acceptance rate, retained code after N minutes (whether accepted code survived), latency percentiles, and cancellation rate.
  • Offline: held-out repos with masked spans, scored on exact match and unit tests passing for generated functions. Agent tasks are scored by tests passing, with cost per task.
  • A/B tests for model or context changes, because small offline gains often don't translate online.

7. Bottlenecks and follow-ups

Likely questionAnswer sketch
"GPU cost too high?"A smaller model, better cache affinity, longer debounce, skipping requests in comments or low-value positions
"Monorepo with 50M lines?"A precomputed symbol index plus embeddings per repo, updated incrementally on commit
"Agent breaks the build?"Run tests in the sandbox before proposing, require review, limit steps and scope
"Latency in distant regions?"More regional pools, and a smaller model there if GPUs are scarce

Key takeaways

  • Inline completion is a latency problem. It needs a small, fast model, a first-token budget of about 150–250 ms, debouncing and cancellation.
  • Completions are prefill-heavy (a few thousand tokens in, dozens out), so prefix caching with session affinity matters enormously.
  • Chat and agent modes use bigger models with repository retrieval (symbols plus embeddings) and sandboxed tool execution.
  • Measure accepted and retained suggestions, not just model benchmarks.

Go deeper