Model Tiers: Frontier, Mid and Small
How LLMs group into tiers by capability, price and speed, how to pick a tier per task, and why most production systems use several models.
The three tiers
Every major provider offers a family of models at different sizes. The names change every few months, but the structure is stable:
| Tier | Typical use | Relative price | Speed |
|---|---|---|---|
| Frontier | Hard reasoning, complex code, agent planning, high-stakes answers | 10–50× small | Slowest |
| Mid | Most chat, RAG answers, summarisation, general agents | 3–10× small | Medium |
| Small / fast | Classification, routing, extraction, simple rewrites, high-volume pipelines | 1× | Fastest |
Open-weight models span the same range, from 1–8B parameter models that run on one GPU or a laptop to hundreds of billions of parameters.
For current prices, see the LLM API cost calculator. This lesson focuses on how to choose.
Matching tiers to tasks
Tasks differ in difficulty far more than traffic suggests:
- Routing and classification ("is this a billing question?"): small models are usually accurate enough, at a fraction of the cost and latency.
- Extraction into a schema from clean text: small or mid.
- Grounded Q&A over retrieved documents: mid. Retrieval does the heavy lifting.
- Multi-step reasoning, ambiguous instructions, long-horizon agents, tricky code: frontier, often with reasoning enabled.
Most mature systems therefore use several models:
Reasoning models and thinking budgets
Many models can think before answering, generating hidden reasoning tokens first. This substantially improves maths, coding, planning and multi-constraint problems, but:
- Thinking tokens are billed as output tokens and can run to thousands per request.
- TTFT grows because the visible answer starts after the thinking.
- Many APIs let you set a thinking budget or effort level. Turn it up for hard tasks and off for simple ones.
Use reasoning where the eval shows a real accuracy gain. It's wasted on "summarise this paragraph".
Other selection criteria
- Context window and output limit: does your largest prompt fit, and can the model write long enough answers?
- Modalities: image, audio or PDF input; image output.
- Tool use and structured output reliability: critical for agents. Test it directly.
- Latency and throughput limits: rate limits per tier, and provisioned-throughput options.
- Data policy: retention, training use, regions and compliance certifications.
- Availability: multiple providers or regions for failover.
Choosing with evals, not leaderboards
Public benchmarks rank general ability. Your task is specific. A practical process:
- Build an eval set of 100–500 real examples with expected outcomes. See Module 7.
- Run candidate models from small to frontier.
- Plot quality against cost per request and pick the cheapest model that clears your quality bar.
- Pin the model version. Re-run the eval before upgrading, because newer is usually better on average but can regress on your cases.
Key takeaways
- Models fall into rough tiers: frontier (best, slowest, priciest), mid (the production workhorse), and small (fast and cheap, for simple tasks).
- Prices between tiers differ 10–50×, so matching the tier to the task is the biggest cost lever you have.
- Reasoning (thinking) modes trade extra output tokens and latency for accuracy on hard problems.
- Pick models with an eval on your own task, not with public leaderboards, and pin versions.