Published 15 July 2026 ยท reference prices, verify before budgeting
Every LLM cost comparison you've seen probably ignores caching. That's a mistake โ because prompt caching changes the cost ranking. Claude charges a write fee; OpenAI doesn't. Gemini's context caching costs per hour stored. Here's the real math.
| Provider | Cache type | Cache write cost | Cache read cost | Cache TTL | Min tokens |
|---|---|---|---|---|---|
| Anthropic (Claude) | Prompt prefix cache | 1.25ร standard input | 0.10ร standard input | 5 min (extendable) | 1,024 |
| OpenAI (GPT-4o) | Automatic cache | No write fee | 0.50ร standard input | ~5 min | 1,024 |
| Google (Gemini) | Context cache | No write fee | 0.25ร standard input | Billed per hour stored | 32,768 |
Key structural differences:
| Model | Input $/1M | Output $/1M |
|---|---|---|
| Claude Sonnet 4 | $3.00 | $15.00 |
| GPT-4o | $2.50 | $10.00 |
| Gemini 2.5 Pro | $1.25 | $10.00 |
Without caching: Gemini 2.5 Pro is cheapest on input ($1.25 vs $2.50 vs $3.00). But once caching applies, the picture changes.
A typical production app: you have a 10,000-token system prompt (instructions + context + examples), and you send 500 requests/day. Each request adds 300 tokens of user input and gets 400 tokens back.
| Model | Input/req (10,300 tok) | Output/req (400 tok) | Monthly (15k req) |
|---|---|---|---|
| GPT-4o | $0.02575 | $0.004 | $446 |
| Claude Sonnet 4 | $0.0309 | $0.006 | $550 |
| Gemini 2.5 Pro | $0.012875 | $0.004 | $255 |
The 10,000-token system prompt is cached. Each subsequent request reads 10,000 cached tokens + 300 new input tokens.
| Model | Cache read /1M | New input /1M | Monthly | vs. no cache |
|---|---|---|---|---|
| GPT-4o (auto cache) | $1.25 | $2.50 | $204 | โ$242 (โ54%) |
| Claude Sonnet 4 + cache | $0.30 | $3.00 | $117 | โ$433 (โ79%) |
| Gemini 2.5 Pro + context cache | $0.3125 | $1.25 | $58* | โ$197 (โ77%) |
*Gemini context cache also charges $4.50/hour for 10k tokens stored. At 1 cache instance running 24/7: +$3.24/month storage cost. Added separately below.
On the first request, Anthropic charges 1.25ร to write the cache. For a 10,000-token system prompt at $3.00/1M: first request costs $0.0375 vs $0.030 standard. The extra $0.0075 is recovered after just 1.14 cache hits. If you're doing 500 requests/day, that write premium is insignificant. For bursty, low-volume use, it adds up.
Gemini's context cache gives 75% off reads โ the biggest discount. But it charges hourly for stored content regardless of traffic. For a 10,000-token cache running 24/7:
For large-context use cases (codebase assistant, document analysis), Gemini's storage cost can erode the cache discount advantage. Calculate both terms before assuming Gemini wins.
| Use case | Best choice | Reason |
|---|---|---|
| High-volume app, large fixed system prompt (>2k tokens, >200 req/day) | Claude Sonnet 4 | 90% cache discount + no storage fee = lowest total at scale |
| Variable traffic, bursty requests | GPT-4o | No write fee, no storage fee, automatic โ most forgiving |
| Very large context (>32k tokens) at high volume | Gemini 2.5 Pro | 75% cache read discount; storage cost small relative to input savings |
| Low volume (<50 req/day), small prompt | Gemini 2.0 Flash | Caching savings minimal; Flash's raw price wins |
| Dev/testing, unpredictable usage | GPT-4o mini | Cheapest at low volume; automatic caching with no overhead |
To compare providers with caching for your use case:
Or just use our cost calculator โ set your token counts and the results reflect standard rates (caching applies automatically on the provider's infrastructure).