Published 15 July 2026 ยท reference prices, verify before budgeting
We took 20 real production workloads โ with realistic token counts for each โ and priced them across four model tiers at 1,000 requests/day. The difference between cheapest and most expensive model for identical tasks reaches 100ร or more.
| Model | Input $/1M | Output $/1M | Tier |
|---|---|---|---|
| Gemini 2.0 Flash cheapest | $0.10 | $0.40 | Small |
| GPT-4o mini | $0.15 | $0.60 | Small |
| Claude Haiku 4.5 | $0.80 | $4.00 | Small+ |
| GPT-4o | $2.50 | $10.00 | Frontier |
| Claude Sonnet 4 | $3.00 | $15.00 | Frontier |
Monthly cost = requests/day ร 30 ร cost per request. Token counts are typical values โ your actual usage may differ. Use the calculator for your exact numbers.
| Workload | Tokens (in/out) | Flash/mo | 4o-mini/mo | Haiku/mo | GPT-4o/mo |
|---|---|---|---|---|---|
| Email classification | 100 / 20 | $0.54 | $0.81 | $4.80 | $18 |
| Sentiment analysis | 150 / 30 | $0.81 | $1.22 | $7.20 | $27 |
| News categorization | 60 / 10 | $0.30 | $0.45 | $2.70 | $10 |
| Search query rewriting | 80 / 60 | $0.95 | $1.44 | $9.00 | $34 |
| Content moderation | 500 / 20 | $1.74 | $2.61 | $15.60 | $58 |
| Workload | Tokens (in/out) | Flash/mo | 4o-mini/mo | Haiku/mo | GPT-4o/mo |
|---|---|---|---|---|---|
| Email reply draft | 300 / 400 | $5.76 | $8.55 | $54 | $202 |
| Product description | 200 / 300 | $4.14 | $6.21 | $39.60 | $148 |
| Social media post | 200 / 120 | $2.02 | $3.06 | $19.20 | $71 |
| FAQ answer | 400 / 350 | $5.40 | $8.10 | $51.60 | $193 |
| User feedback summary | 600 / 200 | $4.20 | $6.30 | $38.40 | $142 |
| Workload | Tokens (in/out) | Flash/mo | 4o-mini/mo | Haiku/mo | Sonnet 4/mo |
|---|---|---|---|---|---|
| Support chat message | 500 / 300 | $5.10 | $7.65 | $48 | $180 |
| Support chat (with history) | 2000 / 300 | $9.60 | $14.40 | $87.60 | $315 |
| Complaint escalation decision | 800 / 100 | $3.60 | $5.40 | $33.60 | $126 |
Context growth is the hidden cost. As conversation history accumulates, input tokens balloon. Message 1 might be 300 tokens; message 10 in the same chat might be 2,500 tokens. That's an 8ร increase in the input cost alone โ see our full analysis of conversation cost growth.
| Workload | Tokens (in/out) | Flash/mo | 4o-mini/mo | Haiku/mo | GPT-4o/mo |
|---|---|---|---|---|---|
| SQL query generation | 400 / 200 | $3.60 | $5.40 | $33.60 | $126 |
| Code review comment | 600 / 800 | $12.36 | $18.54 | $117.60 | $441 |
| Bug fix explanation | 800 / 1000 | $14.40 | $21.60 | $138 | $516 |
| Unit test generation | 600 / 1200 | $16.20 | $24.30 | $166.80 | $625 |
| Workload | Tokens (in/out) | Flash/mo | 4o-mini/mo | Haiku/mo | GPT-4o/mo |
|---|---|---|---|---|---|
| RAG answer (3 chunks) | 2000 / 400 | $10.80 | $16.20 | $102 | $381 |
| Document summarization | 4000 / 500 | $18 | $27 | $156 | $585 |
| Legal clause extraction | 8000 / 1000 | $35.60 | $54 | $336 | $1,260 |
| Invoice data extraction | 1000 / 300 | $6.60 | $9.90 | $58.80 | $220 |
At large context sizes, input cost becomes the dominant term. With a legal document at 8,000 input tokens, Gemini Flash costs $35.60/month at 1k req/day โ GPT-4o costs $1,260. For extraction tasks (structured output, well-defined schema), Flash and 4o-mini typically perform well. For open-ended reasoning over long docs, test quality carefully.
Across all 20 workloads, switching from GPT-4o to Gemini Flash for simple classification saves 95โ97%. On code generation with long output (unit tests, bug fixes), the saving is still 97%. The absolute dollar amounts vary โ $0.30/month for news categorization vs $625/month for unit tests โ but the multiplier is consistent.
The practical implication: don't use one model for everything. Run frontier models on tasks where quality objectively matters (measured by your eval), and use small models for routing, classification, and extraction where benchmarks show they match.
| Task type | Recommended tier | Reasoning |
|---|---|---|
| Classification / routing | Flash or 4o-mini | Output quality meets threshold at 5% of the cost |
| Short generation (<200 tokens) | Flash or 4o-mini | Quality usually acceptable; test first |
| Support chat | Flash or Haiku | Context costs grow fast; Flash wins on volume |
| SQL / code (well-defined) | 4o-mini or Haiku | Better than Flash for structured output; much cheaper than frontier |
| Complex reasoning / code | GPT-4o or Sonnet 4 | Quality gap is real and measurable; budget accordingly |
| Long-context extraction | Flash or 4o-mini | Schema-driven tasks don't need frontier reasoning |
| Long-context analysis | GPT-4o or Gemini 2.5 Pro | Open-ended reasoning over 8k+ tokens needs frontier capability |
All of the above assumed 1,000 requests/day. Scale linearly: 10,000 req/day = 10ร the costs; 100 req/day = 10% of the costs. Use our LLM cost calculator or compare specific models at LLM price comparison.