| Model | Input /MTok | Output /MTok | Provider | Cache read | Notes | Source · verified |
|---|---|---|---|---|---|---|
| Claude Fable 5 | $10 | $50 | Anthropic | $1 | 1M context at standard pricing. Newer tokenizer produces ~30% more tokens for the same text than pre-4.7 Claude models. | source · 2026-09-07 |
| Claude Mythos 5 | $10 | $50 | Anthropic | $1 | Limited availability. | source · 2026-09-07 |
| Claude Opus 5 | $5 | $25 | Anthropic | $0.5 | Fast mode available at $10/$50 per MTok. | source · 2026-09-07 |
| Claude Opus 4.8 | $5 | $25 | Anthropic | $0.5 | Fast mode available at $10/$50 per MTok. 1M context at standard pricing. | source · 2026-09-07 |
| Claude Opus 4.7 | $5 | $25 | Anthropic | $0.5 | Fast mode ($30/$150) deprecated — removed 2026-07-24. | source · 2026-09-07 |
| Claude Opus 4.6 | $5 | $25 | Anthropic | $0.5 | source · 2026-09-07 | |
| Claude Opus 4.5 | $5 | $25 | Anthropic | $0.5 | source · 2026-09-07 | |
| Claude Sonnet 4.6 | $3 | $15 | Anthropic | $0.3 | source · 2026-09-07 | |
| Claude Sonnet 4.5 | $3 | $15 | Anthropic | $0.3 | source · 2026-09-07 | |
| Claude Sonnet 5 | $2 → $3/$15 from 2026-09-01 | $10 | Anthropic | $0.2 | Introductory pricing through 2026-08-31; rises to $3 in / $15 out on 2026-09-01. Newer tokenizer (~+30% tokens). | source · 2026-09-07 |
| Claude Haiku 4.5 | $1 | $5 | Anthropic | $0.1 | source · 2026-09-07 | |
| Gemini 3.1 Pro (preview) | $2 | $12 | — | Prompts >200k tokens: $4 in / $18 out. | source · 2026-09-07 | |
| Gemini 3.5 Flash | $1.5 | $9 | — | source · 2026-09-07 | ||
| Gemini 2.5 Pro | $1.25 | $10 | — | Prompts >200k tokens: $2.50 in / $15 out. | source · 2026-09-07 | |
| Gemini 3.6 Flash | $0.75 | $3.75 | — | source · 2026-09-07 | ||
| Gemini 3 Flash (preview) | $0.5 | $3 | — | Audio input $1.00/MTok. | source · 2026-09-07 | |
| Gemini 3.5 Flash-Lite | $0.3 | $2.5 | — | Same rate for text/image/video/audio input. | source · 2026-09-07 | |
| Gemini 2.5 Flash | $0.3 | $2.5 | — | Audio input $1.00/MTok. | source · 2026-09-07 | |
| Gemini 3.1 Flash-Lite | $0.25 | $1.5 | — | Audio input $0.50/MTok. | source · 2026-09-07 | |
| Gemini 2.5 Flash-Lite | $0.1 | $0.4 | — | Audio input $0.30/MTok. | source · 2026-09-07 | |
| GPT-5.5 Pro | $30 | $180 | OpenAI | — | source · 2026-09-07 | |
| GPT-5.4 Pro | $30 | $180 | OpenAI | — | source · 2026-09-07 | |
| GPT-5.5 | $5 | $30 | OpenAI | — | Named migration target for gpt-5, o3, o1, gpt-4 retirements. | source · 2026-09-07 |
| GPT-5.6 Sol | $4 | $20 | OpenAI | — | Batch pricing 50% off (input and output). | source · 2026-09-07 |
| GPT-5.4 | $2.5 | $15 | OpenAI | — | source · 2026-09-07 | |
| GPT-5.6 Terra | $2 | $12 | OpenAI | — | source · 2026-09-07 | |
| GPT-5.3 Codex | $1.75 | $14 | OpenAI | — | Coding-tuned. | source · 2026-09-07 |
| GPT-5.4 mini | $0.75 | $4.5 | OpenAI | — | Migration target for gpt-3.5-turbo and gpt-5-mini. | source · 2026-09-07 |
| GPT-5.6 Luna | $0.2 | $1.2 | OpenAI | — | source · 2026-09-07 | |
| GPT-5.4 nano | $0.2 | $1.25 | OpenAI | — | source · 2026-09-07 | |
| Claude Opus 4.1 retired | $15 | $75 | Anthropic | $1.5 | Retired on the Claude API (still on Bedrock + Google Cloud). | source · 2026-09-07 |
| Claude Opus 4 retired | $15 | $75 | Anthropic | — | Retired on the Claude API (still on Google Cloud). | source · 2026-09-07 |
| Claude Sonnet 4 retired | $3 | $15 | Anthropic | — | Retired on the Claude API (still on Bedrock + Google Cloud). | source · 2026-09-07 |
| Claude Haiku 3.5 retired | $0.8 | $4 | Anthropic | — | Retired on the Claude API (still on Bedrock + Google Cloud). | source · 2026-09-07 |
| GPT-5 (2025-08-07) deprecated | — | — | OpenAI | — | Deprecated 2026-06-11; shutdown 2026-12-11 → migrate to GPT-5.5. | source · 2026-09-07 |
| o3 (2025-04-16) deprecated | — | — | OpenAI | — | Deprecated 2026-06-11; shutdown 2026-12-11 → migrate to GPT-5.5. | source · 2026-09-07 |
| o1 deprecated | — | — | OpenAI | — | Shutdown 2026-10-23 → migrate to GPT-5.5. | source · 2026-09-07 |
| GPT-4 deprecated | — | — | OpenAI | — | Shutdown 2026-10-23 → migrate to GPT-5.5. | source · 2026-09-07 |
| GPT-3.5 Turbo deprecated | — | — | OpenAI | — | Shutdown 2026-10-23 → migrate to GPT-5.4 mini. | source · 2026-09-07 |
Rows without prices are deprecated/retired models kept for their shutdown dates — see the deprecation tracker. Missing a model you need? We only publish prices we've checked against a provider's own page — request a model.
New to per-token pricing? Start with how to estimate LLM API costs — the cost formula plus the four corrections (tokenizer differences, invisible tokens, caching, deprecations) that turn these rate cards into a real budget.
Frequently asked questions
How often are these prices checked?
Every price is read directly from the provider's official pricing page and stamped with the date we last checked it — currently 2026-09-07. If the data goes more than 30 days without re-verification, every page in this section shows an overdue warning banner instead of quietly going stale.
What does per-MTok pricing mean?
Per million tokens. A model listed at $2.50 input / $15 output charges $2.50 for every 1,000,000 tokens you send it and $15 for every 1,000,000 tokens it generates. A token is roughly 4 characters of English text.
Which LLM API is cheapest right now?
By combined rate, Gemini 2.5 Flash-Lite is the cheapest active model we track, at $0.1 input / $0.4 output per MTok (verified 2026-09-07). Cheapest per token is not always cheapest per task — tokenizers differ between model families.
Do these prices include caching or batch discounts?
No — the table shows base per-token rates. Cache reads (about 0.1x the input price on Anthropic models; see the cache-read column) and batch APIs (about 50% off) can bring real workload costs far below the base rate.