Free tool
LLM cost calculator
Enter your monthly request volume and average token lengths to see the real cost difference between models — cache hit rate included.
Your usage profile
Share of repeated system prompt and context served from cache.
Prices are USD per million tokens, taken from each provider's own pricing page. Indicative only — not a quote. Per-provider cache discounts are included; cache-write cost and promotional or context-tier differences are not.
Price data: 2026-08-18
Estimated monthly cost
Mistral Small 4cheapest
Mistral · $0.15/$0.6 · $0.015
$127
Monthly · $0.0006 per request
GPT-5.6 Luna
OpenAI · $0.2/$1.2 · $0.02
$218
Monthly · $0.0011 per request · 1.7× vs. cheapest
DeepSeek V4 Flash
DeepSeek · $0.44/$1.32 · $0.014
$306
Monthly · $0.0015 per request · 2.4× vs. cheapest
Peak-hour rate; a cache hit takes input cost close to zero.
Mistral Large 3
Mistral · $0.5/$1.5 · $0.05
$364
Monthly · $0.0018 per request · 2.9× vs. cheapest
Gemini 3.5 Flash-Lite
Google · $0.3/$2.5 · $0.03
$410
Monthly · $0.0021 per request · 3.2× vs. cheapest
Gemini 3.7 Flash
Google · $0.75/$3.75 · $0.075
$726
Monthly · $0.0036 per request · 5.7× vs. cheapest
Introductory pricing through 31 Dec 2026; it doubles afterwards.
Grok 4.3
xAI · $1.25/$2.5 · $0.2
$796
Monthly · $0.004 per request · 6.3× vs. cheapest
GPT-5.4 mini
OpenAI · $0.75/$4.5 · $0.075
$816
Monthly · $0.0041 per request · 6.4× vs. cheapest
DeepSeek V4 Pro
DeepSeek · $1.32/$3.96 · $0.044
$919
Monthly · $0.0046 per request · 7.2× vs. cheapest
Peak-hour rate; off-peak hours are charged at half.
Claude Haiku 4.5
Anthropic · $1/$5 · $0.1
$968
Monthly · $0.0048 per request · 7.6× vs. cheapest
Mistral Medium 3.5
Mistral · $1.5/$7.5 · $0.15
$1,452
Monthly · $0.0073 per request · 11.4× vs. cheapest
Grok 4.6
xAI · $2/$6 · $0.5
$1,600
Monthly · $0.008 per request · 12.6× vs. cheapest
Pricing doubles above a 200k-token prompt.
Gemini 3.5 Flash
Google · $1.5/$9 · $0.15
$1,632
Monthly · $0.0082 per request · 12.8× vs. cheapest
GPT-5.6 Terra
OpenAI · $2/$12 · $0.2
$2,176
Monthly · $0.0109 per request · 17.1× vs. cheapest
Gemini 3.1 Pro
Google · $2/$12 · $0.2
$2,176
Monthly · $0.0109 per request · 17.1× vs. cheapest
Input and output rates rise above a 200k-token prompt.
Claude Sonnet 5
Anthropic · $3/$15 · $0.3
$2,904
Monthly · $0.0145 per request · 22.8× vs. cheapest
The best speed/intelligence balance; near-Opus on coding and agentic work.
Claude Opus 5
Anthropic · $5/$25 · $0.5
$4,840
Monthly · $0.0242 per request · 38.1× vs. cheapest
The default for complex agentic coding and enterprise work.
GPT-5.6 Sol
OpenAI · $5/$30 · $0.5
$5,440
Monthly · $0.0272 per request · 42.8× vs. cheapest
Claude Fable 5
Anthropic · $10/$50 · $1
$9,680
Monthly · $0.0484 per request · 76.1× vs. cheapest
For the hardest reasoning and long-horizon agentic work.
GPT-5.5 Pro
OpenAI · $30/$180 · no cache discount
$45,600
Monthly · $0.228 per request · 358.5× vs. cheapest
No cache discount — cost scales fast on long-context workloads.
Let's bring the cost down with architecture
Model choice is only part of the cost. Cache strategy, context length and routing architecture usually make a bigger difference.