GenAI System Design
1. LLM Fundamentals for System Designers

Sampling, Temperature and Determinism

How LLMs choose each token, what temperature and top-p do, why the same prompt can give different answers, and how to design around non-determinism.

Lesson 4 of 6 9 min

From scores to a choice

At each decode step, the model produces a logit (a raw score) for every token in its vocabulary, then turns them into probabilities with softmax. The sampler picks one token from that distribution, appends it, and the next step begins.

Temperature and top-p

The model scores every possible next token. Sampling settings decide how those scores become a choice.

Our API returned a 429 error, which means the client was▍

  • rate67%
  • sending14%
  • throttled11%
  • making4%
  • blocked2%
  • too1%
  • hacked0.20%
  • banana0.01%

Low temperature sharpens the distribution toward the top token, and 0 always picks it. High temperature flattens it, so unlikely tokens like "hacked" start to appear. Top-p cuts the long tail off entirely before sampling. For extraction, classification and JSON, use a low temperature. For brainstorming, go higher.

The knobs

SettingWhat it doesTypical use
TemperatureDivides logits before softmax. Below 1 sharpens toward the top token, above 1 flattens. 0 means always take the top token (greedy)0–0.3 for extraction and classification. 0.7–1.0 for chat and writing
Top-p (nucleus)Keeps the smallest set of tokens whose probabilities sum to p, then samples among them0.9–1.0. Cuts off the nonsense tail
Top-kKeeps only the k most likely tokensSimilar role to top-p. Less common in APIs
Max tokensHard cap on output lengthAlways set it. It controls cost and latency
Stop sequencesEnd generation when a string appearsStructured formats, few-shot templates
SeedFixes the random choices where supportedReproducibility in tests (best effort)

Reasoning models often fix or restrict these settings. Check your provider's docs.

Why "same prompt, different answer"

  1. Sampling. At temperature above 0 the model is intentionally random.
  2. Even at temperature 0, outputs can vary between calls. The server batches your request with others, and different batch shapes change the order of floating-point operations. Tiny numeric differences can flip a close choice between two tokens, and everything after that diverges.
  3. Model updates. A provider's model alias can point to a new snapshot. Pin explicit model versions in production.

Designing around non-determinism

Constrain the output format. Use structured outputs (JSON schema or constrained decoding) so every response parses. Variability in wording is fine. Variability in structure breaks systems.

Validate and retry. Check outputs against the schema and business rules, such as an amount within range or an ID that exists. On failure, retry, preferably with the validation error added to the prompt.

Evaluate over samples, not single runs. A prompt change that "fixed" one example may be noise. Run your eval set (and, for key cases, several samples per case) and compare aggregate scores.

Use self-consistency for hard reasoning. Sample several answers and take the majority, or have a judge pick the best. This costs N× more, so reserve it for high-value decisions.

Cache deliberately. Caching responses gives consistency and savings for identical requests. But a cache also locks in any bad answer, so give entries a TTL and a way to invalidate them.

Don't compare raw outputs in tests. Assert properties instead: it contains the required fields, it cites a source, a judge scores it at least 4 out of 5.

Key takeaways

  • The model outputs a probability for every token in its vocabulary. Sampling settings decide how one is picked.
  • Temperature sharpens (low) or flattens (high) the distribution. Top-p keeps only the most likely tokens covering probability p.
  • Even temperature 0 is not fully deterministic in production, because batching and floating-point effects vary.
  • Design for variability with structured outputs, validation and retries, evals over many samples, and never relying on byte-identical outputs.

Go deeper