LLM API Cost Calculator

Compare what the same workload costs on GPT, Claude, and Gemini: per request, per day, and per month, with prompt caching included.

Prompt + history + retrieved docs

Includes reasoning tokens

%

Repeated prefix hitting the prompt cache

20 rows
Cost for 2K input and 500 output tokens per request at 1000 requests per day
GPT-6 LunaCheapest
OpenAI
$0.00045$0.45$13.501.0×
GPT-5 mini
OpenAI
$0.0015$1.50$45.003.3×
Gemini 3.5 Flash-Lite
Google
$0.0019$1.85$55.504.1×
Gemini 2.5 Flash
Google
$0.0019$1.85$55.504.1×
Gemini 3.8 Flash
Google
$0.0034$3.38$101.257.5×
GPT-5.4 mini
OpenAI
$0.0038$3.75$112.508.3×
Claude Haiku 4.5
Anthropic
$0.0045$4.50$135.0010.0×
Gemini 2.5 Pro
Google
$0.0075$7.50$225.0016.7×
Gemini 3.5 Flash
Google
$0.0075$7.50$225.0016.7×
GPT-4.1
OpenAI
$0.0080$8.00$240.0017.8×
Claude Sonnet 5
Anthropic
$0.0090$9.00$270.0020.0×
GPT-6 Sol
OpenAI
$0.0090$9.00$270.0020.0×
Gemini 3.1 Pro (preview)
Google
$0.01$10.00$300.0022.2×
GPT-5.4
OpenAI
$0.01$12.50$375.0027.8×
Claude Sonnet 4.6
Anthropic
$0.01$13.50$405.0030.0×
Claude Opus 5.5
Anthropic
$0.02$18.00$540.0040.0×
Claude Opus 5
Anthropic
$0.02$22.50$675.0050.0×
GPT-5.5
OpenAI
$0.03$25.00$750.0055.6×
Claude Fable 5.1
Anthropic
$0.04$45.00$1,350100.0×
GPT-6 Astra
OpenAI
$0.04$45.00$1,350100.0×

Standard list prices, no batch discount, 30-day month. “Cheapest” and “vs cheapest” compare against all models, whatever the filters. Anthropic cache writes (1.25× input) aren't included here; the agent calculator models them. Prices last checked 2026-09-25.

20 rows
LLM API prices per 1 million tokens
GPT-6 Luna
OpenAI · Prompts over 272K input tokens: 2x input, 1.5x output
$0.1$0.01$0.51.05M128K
GPT-5 mini
OpenAI
$0.25$0.025$2400K128K
Gemini 3.5 Flash-Lite
Google
$0.3$0.03$2.51.05M65.5K
Gemini 2.5 Flash
Google
$0.3$0.03$2.51.05M65.5K
GPT-5.4 mini
OpenAI
$0.75$0.075$4.5400K128K
Gemini 3.8 Flash
Google · Promotional price through Dec 31, 2026 ($1.50 / $7.50 after)
$0.75$0.075$3.751.05M65.5K
Claude Haiku 4.5
Anthropic
$1$0.1$5200K64K
Gemini 2.5 Pro
Google · Prompts over 200K tokens: $2.50 / $15
$1.25$0.125$101.05M65.5K
Gemini 3.5 Flash
Google
$1.5$0.15$91.05M65.5K
Claude Sonnet 5
Anthropic
$2$0.2$101M128K
GPT-6 Sol
OpenAI · Prompts over 272K input tokens: 2x input, 1.5x output
$2$0.2$101.05M128K
GPT-4.1
OpenAI
$2$0.5$81.05M32.8K
Gemini 3.1 Pro (preview)
Google · Prompts over 200K tokens: $4 / $18
$2$0.2$121.05M65.5K
GPT-5.4
OpenAI · Prompts over 272K input tokens: 2x input, 1.5x output
$2.5$0.25$151.05M128K
Claude Sonnet 4.6
Anthropic
$3$0.3$151M128K
Claude Opus 5.5
Anthropic
$4$0.2$201M128K
Claude Opus 5
Anthropic
$5$0.5$251M128K
GPT-5.5
OpenAI · Prompts over 272K input tokens: 2x input, 1.5x output
$5$0.5$301.05M128K
Claude Fable 5.1
Anthropic
$10$0.25$501M128K
GPT-6 Astra
OpenAI · Prompts over 272K input tokens: 2x input, 1.5x output
$10$1$501.05M128K

USD per 1M tokens, standard tier, short-context rates. Sources: official Anthropic, OpenAI, and Google pricing pages, last checked 2026-09-25.

How to estimate your LLM API bill

  1. Measure real token counts. Log usage from a sample of real requests, or paste representative prompts into the token counter. Include the system prompt, tool schemas, conversation history, and retrieved context.
  2. Separate input from output. Output costs several times more than input, so a chatty model or high reasoning effort can dominate the bill.
  3. Estimate your cache-hit share. A stable system prompt and tool list in front of a short user message can mean 70–90% of input tokens are cache hits.
  4. Multiply by volume. Requests per day × 30. For agents, one user task can mean 10–50 model calls, so use the agent cost calculator.

Ways to cut LLM API costs

  • Prompt caching: put static content first and variable content last so the prefix can be reused.
  • Batch APIs: Anthropic and OpenAI give 50% off for asynchronous jobs that can wait up to 24 hours.
  • Model routing: send classification, extraction, and sub-agent work to a small model, and keep the frontier model for planning and hard reasoning.
  • Shorter outputs: ask for structured, concise answers and tune reasoning effort per route.

Long-context pricing

Some models charge more once a prompt passes a threshold. Gemini Pro models bill prompts over 200K tokens at higher rates, and OpenAI's GPT-5.4 and later charge 2× input and 1.5× output above 272K input tokens. Claude 4.6 and later bill the full 1M context window at standard rates. The table uses short-context prices; check the notes in the pricing tab for each model's long-context rates.

Frequently asked questions

How is LLM API cost calculated?

Cost per request = input tokens × input price + output tokens × output price, with prices quoted per million tokens. Any input served from the prompt cache is billed at the lower cached-input rate instead. Multiply by requests per day and by 30 for a monthly estimate.

Why are output tokens more expensive than input tokens?

Input tokens are processed in parallel in one forward pass, but output tokens are generated one at a time, each needing its own pass through the model. That makes output 4–8× more expensive at most providers. Reasoning or thinking tokens are billed as output.

What is cached input pricing?

When the start of a prompt (system prompt, tool definitions, long documents) is the same as a recent request, providers can reuse the cached computation and charge a fraction of the input price, typically 10% at Anthropic, OpenAI's newer models, and Gemini. Anthropic also charges a one-time 1.25× premium to write to the 5-minute cache.

Which LLM API is the cheapest?

It depends on your input-to-output ratio and caching. Small models like GPT-6 Luna, Gemini Flash-Lite, and GPT-5 mini cost cents per million tokens, but for agentic work the model that finishes a task in the fewest steps is often cheaper overall. Compare cost per completed task, not per token.

How often are these prices updated?

Prices come from the official Anthropic, OpenAI, and Google pricing pages and were last checked on 2026-09-25. Always confirm on the provider's page before committing to a budget.

Related reading

More AI builder tools

  • LLM Token Counter — Count tokens for GPT, Claude, and Gemini, check context-window fit, and see what the prompt costs.
  • LLM VRAM Calculator — Estimate GPU memory for weights, KV cache, and fine-tuning, and see which GPUs a model fits on.
  • AI Agent Cost Calculator — Model the real cost of agent loops: growing context, tool results, prompt caching, retries, and sub-agents.