Tokens and Tokenizers
What a token is, why it is the unit of cost, latency and limits in every LLM system, and how tokenization quirks affect design.
The unit everything is measured in
An LLM never sees characters or words. Before text reaches the model, a tokenizer splits it into tokens, which are pieces from a fixed vocabulary, and maps each one to an integer ID. The model reads a sequence of IDs and predicts the next ID. The tokenizer turns output IDs back into text.
For a system designer, tokens are the unit of almost everything:
| What | Measured in |
|---|---|
| Price | $ per million input tokens and per million output tokens |
| Context window | Max tokens of prompt plus output |
| Speed | Tokens per second (decode), time to first token |
| Rate limits | Tokens per minute (TPM) as well as requests per minute |
| GPU memory | KV cache bytes per token |
How tokenizers split text
Modern tokenizers use byte-pair encoding (BPE) or similar schemes. They start from bytes and repeatedly merge the most frequent pairs seen in training data. The result:
- Common words are one token: "the", " model", " request".
- Rare words are split into pieces: "Kubernetes" might be "Kub" + "ernetes".
- Spaces usually attach to the following word: " apple" is a different token from "apple".
- Numbers and IDs are split into chunks of digits.
See how text becomes tokens
Uses the o200k_base tokenizer (GPT-4o and later). Other providers split text differently, but the patterns are the same.
Try the Hindi and Japanese samples: the same meaning can cost 2–4× more tokens than English. Long numbers and IDs split into many pieces, which is one reason models make arithmetic and copying mistakes.
Rules of thumb
Design implications
1. Cost and latency depend on the language and format. A product serving Hindi or Japanese users may pay 2× more per conversation than for English. JSON with long keys and whitespace wastes tokens compared with compact formats. Budget per locale, and test your real traffic.
2. Count with the right tokenizer. Each model family has its own vocabulary. The same text might be 1,000 tokens for one provider and 1,150 for another. OpenAI publishes its tokenizers (tiktoken). Others expose a token-counting API endpoint. When enforcing a context limit, count with the target model's tokenizer or leave a 10–20% margin.
3. Rate limits are in tokens. Providers limit tokens per minute as well as requests. One user pasting a 100K-token document can use up a whole minute's budget. Your gateway should meter tokens, not just requests. Module 6 covers token-based rate limiting.
4. Tokenization explains some model failures. Models see "4481-2293" as a few arbitrary chunks, not digits, so exact arithmetic, character counting and copying long IDs are error-prone. Don't ask the model to count characters or do arithmetic on large numbers. Give it a tool (a calculator or code execution) instead.
5. Special tokens structure conversations. Chat APIs wrap messages in special tokens that mark the system, user and assistant roles. You don't see them, but they count toward the context. Tool definitions and JSON schemas also add hidden prompt tokens, often hundreds or thousands per request.
Key takeaways
- Models read and write tokens, which are sub-word pieces from a fixed vocabulary of about 100K–260K entries, not characters or words.
- English prose averages about 4 characters or 0.75 words per token. Code, numbers and many non-English languages use more tokens for the same content.
- Every limit and price is in tokens, including context windows, rate limits, per-token pricing and generation speed.
- Token counts differ between model families, so count with the right tokenizer or budget a margin.