GenAI System Design
6. Integrating LLMs into Existing Systems

Batch and Async LLM Jobs

How to run LLM work that doesn't need an instant answer, using queues, workers, provider batch APIs, idempotency, partial failures and backfills, and why it cuts cost roughly in half.

Lesson 5 of 6 8 min

Not everything needs to be real-time

Look at a typical AI product's LLM usage and much of it doesn't block a user:

  • Summarising every support ticket when it closes
  • Classifying and tagging product reviews
  • Generating embeddings for new documents
  • Enriching CRM records, extracting fields from invoices
  • Nightly eval runs, and backfills after a prompt change

Running these synchronously wastes money and makes them fragile. Run them asynchronously.

Two async patterns

1. Your own queue and workers. Minutes-level latency, and full control.

Producer
event or scheduler
Queue
SQS, Kafka, Pub/Sub
Workers
concurrency tied to TPM budget
LLM
via gateway
Results store
status per item

2. Provider batch APIs. Upload a file of many requests, and the provider processes it within a window (often up to 24 hours) at a discount of around 50%, with separate, higher rate limits. They suit large, latency-tolerant jobs such as backfills, eval sweeps and bulk enrichment.

Sync APIOwn queue + sync APIProvider batch API
LatencySecondsSeconds to minutesMinutes to hours
PriceListList≈ 50% off (typical)
Rate limitsShared with interactiveShared: throttle itSeparate batch quota
ControlFullFullSubmit and poll

For self-hosted models, batch work runs at large batch sizes on the same or dedicated GPUs, often in off-peak hours, which gives much higher throughput per GPU (see batching).

Worker design

  • Concurrency from the token budget: workers × average tokens per request must stay under your TPM allocation. Otherwise the batch job causes 429s for interactive users.
  • Lower priority: batch traffic goes through the gateway in a low-priority lane.
  • Idempotency: key each item by (job ID, item ID, prompt version) and skip already-completed items on rerun.
  • Per-item retries with backoff, then a dead-letter queue for items that keep failing (content policy refusals, malformed input).
  • Checkpoint progress so a crash mid-job resumes rather than restarts.

Partial failures and validation

At 1M items, some will fail or return bad output:

  • Validate every output against its schema and business rules, and mark invalid ones for retry, possibly with a stronger model.
  • Track status per item: pending, done, failed, invalid. Report counts and a sample of failures.
  • Version your prompts, so you know which prompt produced each result and can backfill selectively after a change.

Estimating a batch job

Example: summarise 2M support tickets, 1,500 input and 150 output tokens each, on a model at $0.25 / $2 per million:

  • Input: 2M × 1,500 = 3B tokens × $0.25/1M = $750
  • Output: 2M × 150 = 300M tokens × $2/1M = $600
  • ≈ $1,350 at list price, ≈ $675 via a batch API
  • Time at a 2M TPM allocation: 3.3B tokens ÷ 2M per minute ≈ 28 hours via your own workers. That's a good reason to use the batch API's separate quota.

Key takeaways

  • Much LLM work isn't interactive, such as summaries, classification, enrichment, embeddings and evals. Run it asynchronously.
  • Provider batch APIs typically cost about 50% less in exchange for results within hours.
  • Use queues and workers with concurrency limits tied to your token budget, idempotent jobs, and per-item retries.
  • Design for partial failure. Track status per item, retry failures separately, and make reruns safe.

Go deeper