LLM API Cost Calculator

Real provider pricing — every number verified against the official pricing pages and stamped 2026-08-06. Pick models, describe your workload, get monthly cost side by side.

Models to compare — tap to toggle

$185/mo

Gemini 2.5 Flash is the cheapest of 3 selected — 100,000 requests × (2,000 in + 500 out tokens) at verified rates.

ModelPer requestMonthlyvs cheapest
Gemini 2.5 Flash$0.00185$185
Claude Sonnet 5$0.00900$900+386%
GPT-5.4$0.01250$1,250+576%

Straight rate-card math — no caching/batch discounts applied. Tokenizers differ between families (Anthropic’s newest ≈+30% tokens for identical text); benchmark with your own payloads.

Quick reference: cost per 1M input + 1M output tokens

ModelInput /MTokOutput /MTok1M in + 1M outProvider
Gemini 2.5 Flash-Lite$0.1$0.4$0.5Google
GPT-5.6 Luna$0.2$1.2$1.4OpenAI
GPT-5.4 nano$0.2$1.25$1.45OpenAI
Gemini 3.1 Flash-Lite$0.25$1.5$1.75Google
Gemini 3.5 Flash-Lite$0.3$2.5$2.8Google
Gemini 2.5 Flash$0.3$2.5$2.8Google
Gemini 3 Flash (preview)$0.5$3$3.5Google
GPT-5.4 mini$0.75$4.5$5.25OpenAI
Claude Haiku 4.5$1$5$6Anthropic
Gemini 3.6 Flash$1.5$7.5$9Google
Gemini 3.5 Flash$1.5$9$10.5Google
Gemini 2.5 Pro$1.25$10$11.25Google
Claude Sonnet 5$2$10$12Anthropic
GPT-5.6 Terra$2$12$14OpenAI
Gemini 3.1 Pro (preview)$2$12$14Google
GPT-5.3 Codex$1.75$14$15.75OpenAI
GPT-5.4$2.5$15$17.5OpenAI
Claude Sonnet 4.6$3$15$18Anthropic
Claude Sonnet 4.5$3$15$18Anthropic
Claude Opus 5$5$25$30Anthropic
Claude Opus 4.8$5$25$30Anthropic
Claude Opus 4.7$5$25$30Anthropic
Claude Opus 4.6$5$25$30Anthropic
Claude Opus 4.5$5$25$30Anthropic
GPT-5.6 Sol$5$30$35OpenAI
GPT-5.5$5$30$35OpenAI
Claude Fable 5$10$50$60Anthropic
Claude Mythos 5$10$50$60Anthropic
GPT-5.5 Pro$30$180$210OpenAI
GPT-5.4 Pro$30$180$210OpenAI

Sorted cheapest-first by combined rate. Full table with cache pricing, notes and sources: pricing table.

Three things pricing tables won't tell you

  • Tokenizers differ. Anthropic's newer models (Fable 5, Sonnet 5, Opus 4.7+) use a tokenizer that produces roughly 30% more tokens for the same text than earlier Claude models — so $/token comparisons across families aren't $/task comparisons. Benchmark with your own payloads.
  • Prices are scheduled to move. Claude Sonnet 5 rises from $2/$10 to $3/$15 per MTok on Sept 1, 2026 — a +50% jump already announced. Details →
  • Caching and batch change everything. Cache reads run ~0.1× input price (Anthropic) and batch APIs run ~50% off — a cache-heavy agent workload can cost a fraction of the naive estimate.

Frequently asked questions

How do I estimate tokens per request?

English prose runs roughly 4 characters per token as a floor; code, JSON and non-English text run heavier. Count the invisible tokens too — system prompts, tool definitions and conversation history all bill as input on every request. Anthropic's newest models (Fable 5, Sonnet 5, Opus 4.7 and later) also produce about 30% more tokens for the same text than earlier Claude models.

How is cached input priced?

Cache reads on Anthropic models run about 0.1x the normal input rate — Claude Sonnet 4.6 charges $0.30 per MTok for cached input against $3 for fresh input (verified 2026-08-06). If most of your requests share a long prefix, caching changes the estimate by an order of magnitude.

Do batch APIs get a discount?

Yes, roughly 50%. OpenAI's batch pricing takes 50% off both input and output, and Anthropic's Batch API halves both sides — Claude Sonnet 5 at post-increase rates becomes $1.50/$7.50 per MTok in batch.

Does this calculator apply those discounts?

No — it is straight rate-card math on verified base prices. Caching and batch savings depend on your workload's shape, so they are called out next to every result rather than silently modeled.