$185/mo
Gemini 2.5 Flash is the cheapest of 3 selected — 100,000 requests × (2,000 in + 500 out tokens) at verified rates.
| Model | Per request | Monthly | vs cheapest |
|---|---|---|---|
| Gemini 2.5 Flash | $0.00185 | $185 | — |
| Claude Sonnet 5 | $0.00900 | $900 | +386% |
| GPT-5.4 | $0.01250 | $1,250 | +576% |
Straight rate-card math — no caching/batch discounts applied. Tokenizers differ between families (Anthropic’s newest ≈+30% tokens for identical text); benchmark with your own payloads.
Quick reference: cost per 1M input + 1M output tokens
| Model | Input /MTok | Output /MTok | 1M in + 1M out | Provider |
|---|---|---|---|---|
| Gemini 2.5 Flash-Lite | $0.1 | $0.4 | $0.5 | |
| GPT-5.6 Luna | $0.2 | $1.2 | $1.4 | OpenAI |
| GPT-5.4 nano | $0.2 | $1.25 | $1.45 | OpenAI |
| Gemini 3.1 Flash-Lite | $0.25 | $1.5 | $1.75 | |
| Gemini 3.5 Flash-Lite | $0.3 | $2.5 | $2.8 | |
| Gemini 2.5 Flash | $0.3 | $2.5 | $2.8 | |
| Gemini 3 Flash (preview) | $0.5 | $3 | $3.5 | |
| GPT-5.4 mini | $0.75 | $4.5 | $5.25 | OpenAI |
| Claude Haiku 4.5 | $1 | $5 | $6 | Anthropic |
| Gemini 3.6 Flash | $1.5 | $7.5 | $9 | |
| Gemini 3.5 Flash | $1.5 | $9 | $10.5 | |
| Gemini 2.5 Pro | $1.25 | $10 | $11.25 | |
| Claude Sonnet 5 | $2 | $10 | $12 | Anthropic |
| GPT-5.6 Terra | $2 | $12 | $14 | OpenAI |
| Gemini 3.1 Pro (preview) | $2 | $12 | $14 | |
| GPT-5.3 Codex | $1.75 | $14 | $15.75 | OpenAI |
| GPT-5.4 | $2.5 | $15 | $17.5 | OpenAI |
| Claude Sonnet 4.6 | $3 | $15 | $18 | Anthropic |
| Claude Sonnet 4.5 | $3 | $15 | $18 | Anthropic |
| Claude Opus 5 | $5 | $25 | $30 | Anthropic |
| Claude Opus 4.8 | $5 | $25 | $30 | Anthropic |
| Claude Opus 4.7 | $5 | $25 | $30 | Anthropic |
| Claude Opus 4.6 | $5 | $25 | $30 | Anthropic |
| Claude Opus 4.5 | $5 | $25 | $30 | Anthropic |
| GPT-5.6 Sol | $5 | $30 | $35 | OpenAI |
| GPT-5.5 | $5 | $30 | $35 | OpenAI |
| Claude Fable 5 | $10 | $50 | $60 | Anthropic |
| Claude Mythos 5 | $10 | $50 | $60 | Anthropic |
| GPT-5.5 Pro | $30 | $180 | $210 | OpenAI |
| GPT-5.4 Pro | $30 | $180 | $210 | OpenAI |
Sorted cheapest-first by combined rate. Full table with cache pricing, notes and sources: pricing table.
Three things pricing tables won't tell you
- Tokenizers differ. Anthropic's newer models (Fable 5, Sonnet 5, Opus 4.7+) use a tokenizer that produces roughly 30% more tokens for the same text than earlier Claude models — so $/token comparisons across families aren't $/task comparisons. Benchmark with your own payloads.
- Prices are scheduled to move. Claude Sonnet 5 rises from $2/$10 to $3/$15 per MTok on Sept 1, 2026 — a +50% jump already announced. Details →
- Caching and batch change everything. Cache reads run ~0.1× input price (Anthropic) and batch APIs run ~50% off — a cache-heavy agent workload can cost a fraction of the naive estimate.
Frequently asked questions
How do I estimate tokens per request?
English prose runs roughly 4 characters per token as a floor; code, JSON and non-English text run heavier. Count the invisible tokens too — system prompts, tool definitions and conversation history all bill as input on every request. Anthropic's newest models (Fable 5, Sonnet 5, Opus 4.7 and later) also produce about 30% more tokens for the same text than earlier Claude models.
How is cached input priced?
Cache reads on Anthropic models run about 0.1x the normal input rate — Claude Sonnet 4.6 charges $0.30 per MTok for cached input against $3 for fresh input (verified 2026-08-06). If most of your requests share a long prefix, caching changes the estimate by an order of magnitude.
Do batch APIs get a discount?
Yes, roughly 50%. OpenAI's batch pricing takes 50% off both input and output, and Anthropic's Batch API halves both sides — Claude Sonnet 5 at post-increase rates becomes $1.50/$7.50 per MTok in batch.
Does this calculator apply those discounts?
No — it is straight rate-card math on verified base prices. Caching and batch savings depend on your workload's shape, so they are called out next to every result rather than silently modeled.