| Model | Input /MTok | Output /MTok | Provider | Cache read | Notes | Source · verified |
|---|---|---|---|---|---|---|
| Claude Fable 5 | $10 | $50 | Anthropic | $1 | 1M context at standard pricing. Newer tokenizer produces ~30% more tokens for the same text than pre-4.7 Claude models. | source · 2026-08-06 |
| Claude Mythos 5 | $10 | $50 | Anthropic | $1 | Limited availability. | source · 2026-08-06 |
| Claude Opus 5 | $5 | $25 | Anthropic | $0.5 | Fast mode available at $10/$50 per MTok. | source · 2026-08-06 |
| Claude Opus 4.8 | $5 | $25 | Anthropic | $0.5 | Fast mode available at $10/$50 per MTok. 1M context at standard pricing. | source · 2026-08-06 |
| Claude Opus 4.7 | $5 | $25 | Anthropic | $0.5 | Fast mode ($30/$150) deprecated — removed 2026-07-24. | source · 2026-08-06 |
| Claude Opus 4.6 | $5 | $25 | Anthropic | $0.5 | source · 2026-08-06 | |
| Claude Opus 4.5 | $5 | $25 | Anthropic | $0.5 | source · 2026-08-06 | |
| Claude Sonnet 4.6 | $3 | $15 | Anthropic | $0.3 | source · 2026-08-06 | |
| Claude Sonnet 4.5 | $3 | $15 | Anthropic | $0.3 | source · 2026-08-06 | |
| Claude Sonnet 5 | $2 → $3/$15 from 2026-09-01 | $10 | Anthropic | $0.2 | Introductory pricing through 2026-08-31; rises to $3 in / $15 out on 2026-09-01. Newer tokenizer (~+30% tokens). | source · 2026-08-06 |
| Claude Haiku 4.5 | $1 | $5 | Anthropic | $0.1 | source · 2026-08-06 | |
| Gemini 3.1 Pro (preview) | $2 | $12 | — | Prompts >200k tokens: $4 in / $18 out. | source · 2026-08-06 | |
| Gemini 3.6 Flash | $1.5 | $7.5 | — | source · 2026-08-06 | ||
| Gemini 3.5 Flash | $1.5 | $9 | — | source · 2026-08-06 | ||
| Gemini 2.5 Pro | $1.25 | $10 | — | Prompts >200k tokens: $2.50 in / $15 out. | source · 2026-08-06 | |
| Gemini 3 Flash (preview) | $0.5 | $3 | — | Audio input $1.00/MTok. | source · 2026-08-06 | |
| Gemini 3.5 Flash-Lite | $0.3 | $2.5 | — | Same rate for text/image/video/audio input. | source · 2026-08-06 | |
| Gemini 2.5 Flash | $0.3 | $2.5 | — | Audio input $1.00/MTok. | source · 2026-08-06 | |
| Gemini 3.1 Flash-Lite | $0.25 | $1.5 | — | Audio input $0.50/MTok. | source · 2026-08-06 | |
| Gemini 2.5 Flash-Lite | $0.1 | $0.4 | — | Audio input $0.30/MTok. | source · 2026-08-06 | |
| GPT-5.5 Pro | $30 | $180 | OpenAI | — | source · 2026-08-06 | |
| GPT-5.4 Pro | $30 | $180 | OpenAI | — | source · 2026-08-06 | |
| GPT-5.6 Sol | $5 | $30 | OpenAI | — | Batch pricing 50% off (input and output). | source · 2026-08-06 |
| GPT-5.5 | $5 | $30 | OpenAI | — | Named migration target for gpt-5, o3, o1, gpt-4 retirements. | source · 2026-08-06 |
| GPT-5.4 | $2.5 | $15 | OpenAI | — | source · 2026-08-06 | |
| GPT-5.6 Terra | $2 | $12 | OpenAI | — | source · 2026-08-06 | |
| GPT-5.3 Codex | $1.75 | $14 | OpenAI | — | Coding-tuned. | source · 2026-08-06 |
| GPT-5.4 mini | $0.75 | $4.5 | OpenAI | — | Migration target for gpt-3.5-turbo and gpt-5-mini. | source · 2026-08-06 |
| GPT-5.6 Luna | $0.2 | $1.2 | OpenAI | — | source · 2026-08-06 | |
| GPT-5.4 nano | $0.2 | $1.25 | OpenAI | — | source · 2026-08-06 | |
| Claude Opus 4.1 retired | $15 | $75 | Anthropic | $1.5 | Retired on the Claude API (still on Bedrock + Google Cloud). | source · 2026-08-06 |
| Claude Opus 4 retired | $15 | $75 | Anthropic | — | Retired on the Claude API (still on Google Cloud). | source · 2026-08-06 |
| Claude Sonnet 4 retired | $3 | $15 | Anthropic | — | Retired on the Claude API (still on Bedrock + Google Cloud). | source · 2026-08-06 |
| Claude Haiku 3.5 retired | $0.8 | $4 | Anthropic | — | Retired on the Claude API (still on Bedrock + Google Cloud). | source · 2026-08-06 |
| GPT-5 (2025-08-07) deprecated | — | — | OpenAI | — | Deprecated 2026-06-11; shutdown 2026-12-11 → migrate to GPT-5.5. | source · 2026-08-06 |
| o3 (2025-04-16) deprecated | — | — | OpenAI | — | Deprecated 2026-06-11; shutdown 2026-12-11 → migrate to GPT-5.5. | source · 2026-08-06 |
| o1 deprecated | — | — | OpenAI | — | Shutdown 2026-10-23 → migrate to GPT-5.5. | source · 2026-08-06 |
| GPT-4 deprecated | — | — | OpenAI | — | Shutdown 2026-10-23 → migrate to GPT-5.5. | source · 2026-08-06 |
| GPT-3.5 Turbo deprecated | — | — | OpenAI | — | Shutdown 2026-10-23 → migrate to GPT-5.4 mini. | source · 2026-08-06 |
Rows without prices are deprecated/retired models kept for their shutdown dates — see the deprecation tracker. Missing a model you need? We only publish prices we've checked against a provider's own page — request a model.
Frequently asked questions
How often are these prices checked?
Every price is read directly from the provider's official pricing page and stamped with the date we last checked it — currently 2026-08-06. If the data goes more than 30 days without re-verification, every page in this section shows an overdue warning banner instead of quietly going stale.
What does per-MTok pricing mean?
Per million tokens. A model listed at $2.50 input / $15 output charges $2.50 for every 1,000,000 tokens you send it and $15 for every 1,000,000 tokens it generates. A token is roughly 4 characters of English text.
Which LLM API is cheapest right now?
By combined rate, Gemini 2.5 Flash-Lite is the cheapest active model we track, at $0.1 input / $0.4 output per MTok (verified 2026-08-06). Cheapest per token is not always cheapest per task — tokenizers differ between model families.
Do these prices include caching or batch discounts?
No — the table shows base per-token rates. Cache reads (about 0.1x the input price on Anthropic models; see the cache-read column) and batch APIs (about 50% off) can bring real workload costs far below the base rate.