HomeBlog › The frontier tax

The frontier tax: dropping one model tier saves 40% or 96% depending on the vendor

Published 2 August 2026 · reference prices from our 387-model registry, verify before budgeting

"Just use the small model" is the standard cost advice. I finally priced it properly across our whole model registry and the advice turns out to be almost meaningless on its own. Downgrading one tier inside Anthropic's lineup saves you 40%. The exact same move inside Google's lineup saves 96%. Same decision, same amount of engineering work, wildly different payoff — because the gap between a vendor's flagship and its cheap tier is not a constant. It ranges from 1.7× to 24×.

The workload I priced everything on

One month, 200M input tokens and 40M output tokens. That's roughly 250,000 requests at 800 tokens in and 160 out — a real-shaped product doing classification, summaries and short chat replies. No caching, no batch discount, so the numbers stay comparable. Every price below comes from our tracked registry (387 models, reference prices for July–August 2026).

What one tier down actually costs you

FamilyTier$/1M in$/1M outMonth
OpenAIGPT-5.5$5.00$30.00$2,200
GPT-5$1.25$10.00$650
GPT-5 mini$0.40$1.60$144
GPT-5 nano cheapest$0.10$0.40$36
AnthropicClaude Opus 4.8$5.00$25.00$2,000
Claude Sonnet 4.6$3.00$15.00$1,200
Claude Haiku 4.5$1.00$5.00$400
GoogleGemini 3.1 Pro$2.00$12.00$880
Gemini 3.1 Flash$0.10$0.40$36
Gemini 3.1 Flash-Lite$0.05$0.20$18
xAIGrok 4$3.00$15.00$1,200
Grok 4 mini$0.50$2.50$200
DeepSeekDeepSeek V4$0.27$1.10$98
DeepSeek V4 Flash$0.14$0.28$39
MistralMistral Large$2.00$6.00$640
Mistral Small$0.10$0.30$32
⚠️ Reference prices, August 2026. Always verify on the provider's page before budgeting. Run your own token mix through the AI API cost calculator. · Report outdated price →

The one-step-down table is the interesting one

Strip out everything except the move most teams actually consider — flagship to the tier immediately below it:

MoveBefore → afterSaved
Opus 4.8 → Sonnet 4.6$2,000 → $1,20040%
DeepSeek V4 → V4 Flash$98 → $3960%
GPT-5.5 → GPT-5$2,200 → $65070%
Grok 4 → Grok 4 mini$1,200 → $20083%
Mistral Large → Small$640 → $3295%
Gemini 3.1 Pro → Flash$880 → $3696%

Forty percent versus ninety-six percent. Both are "use the smaller model." One of them is worth a sprint of eval work and a routing layer; the other is worth an afternoon at best, because you'll spend more engineer-hours proving Sonnet is good enough than the $800 you save.

Why Anthropic's gap is so narrow

This isn't Anthropic being stingy with discounts. It's the opposite: their flagship got dramatically cheaper. Claude Opus 4.1 sits in our registry at $15/$75. Opus 4.5, 4.6 and 4.8 all sit at $5/$25 — a 3× cut at the top of the lineup with no change to the Sonnet or Haiku price. On my 240M-token month that took Opus from $6,000 to $2,000.

The side effect: the Opus-to-Sonnet gap collapsed from 5× to 1.67×. The frontier tax at Anthropic mostly stopped existing. OpenAI did the same thing on a different axis — o1 at $15/$60 was replaced by GPT-5 at $1.25/$10, so the expensive-frontier-model line item people budgeted for in 2025 is a different order of magnitude now.

Google went the other way. They kept Pro reasonably priced at $2/$12 and pushed Flash-Lite down to $0.05/$0.20, so their internal spread is now the widest of any major family. Nothing about that is accidental — it's a routing pitch.

The cross-vendor number that annoyed me

Gemini 3.1 Flash costs $36 on this workload. DeepSeek V4, which I'd mentally filed as "the cheap option," costs $98 — nearly 3× more. And GPT-5.5 at $2,200 is 122× the price of Gemini 3.1 Flash-Lite at $18 for moving the same token volume.

I'd been carrying a two-year-old mental price ranking. It was wrong in both directions. Cheap-vendor reputation lags actual price sheets by about a year, and the numbers move faster than the folklore.

What I do with this

1. I check the tier gap before I build the router. If flagship-to-mini is under 2×, a cascade isn't worth the complexity or the eval budget — I take the flagship and spend the effort on prompt caching instead, which cuts more.

2. If the gap is 20×, I route aggressively. Cheap tier handles classification, extraction, tagging and routing. The flagship sees only what fails a confidence check. On the Google spread that turns $880 into something close to $60 with a 10% escalation rate — the arithmetic is in the LLM cascade savings calculator and the model routing savings calculator.

3. I re-price the whole lineup every quarter. The Opus repricing changed which optimizations were worth doing, and I found out weeks late. That's what the price-change tracker is for now.

4. I never compare on input price alone. Output runs 4× input at the median across the 97 first-party chat models I checked, and 6× on GPT-5.5 and Gemini 3.1 Pro. A chatty model on a cheap input rate loses to a terse one on an expensive rate more often than people expect. If your app is switching models, the migration cost matters too — see the model migration cost calculator.

Put your own token split into the AI API Cost Calculator →, or model a full product in the AI app cost estimator and the chatbot cost calculator. New to token pricing? Start at Learn: how API pricing works.

Reference estimates, August 2026, taken from our tracked model registry. Prices change without notice and exclude caching, batch and volume discounts — verify with the provider before you commit a budget. Not affiliated with OpenAI, Anthropic, Google, xAI, DeepSeek or Mistral.