The frontier price collapse, visualized
Input price per 1M tokens for the best generally-available "frontier-class" model at each date. Same tier of capability, three years apart.
Not the same model — the same class of capability. That distinction is the entire story: models don't get cheaper, capability tiers do, as newer models make yesterday's frontier a commodity.
Timeline of every major price event
The price that defined "expensive AI." A single heavy chat session could cost dollars. GPT-3.5 Turbo, at $2/$2, was the volume workhorse.
Anthropic undercuts GPT-4 by ~60% with a 100k context window — the first real price pressure at the frontier tier.
OpenAI's first big cut. Faster, 128k context, and a third of the price. The pattern is set: each generation ships better and cheaper.
The three-tier menu becomes the industry template. Haiku at $0.25 makes million-token workloads casual; Opus stakes out the premium ceiling.
Near-frontier quality at two orders of magnitude below GPT-4's launch price. Triggers an immediate price war among Chinese providers (Alibaba, ByteDance cut within days) and permanently resets what "cheap" means.
Half of Turbo's price, multimodal included. Frontier input is now 6× cheaper than 14 months earlier.
Better than GPT-3.5 Turbo at less than a third of its price. The "budget tier" now outperforms 2023's frontier on most tasks.
Second halving in four months. Frontier input has fallen 12× since GPT-4 launch — 17 months.
Gemini 1.5 Flash drops from $0.35 to $0.075 input; 1.5 Pro from $3.50 to $1.25. Google buys market share with the steepest sustained cuts of any major provider.
The counter-example worth remembering: Anthropic priced the new Haiku 3.2× above its predecessor, arguing capability justified it. Deflation is a trend, not a law — budget tiers can and do reprice upward when a model overdelivers.
Reasoning models reset the ceiling — and introduce invisible "thinking token" costs, where you pay for chain-of-thought you never see. A new premium tier is born at 2023 prices.
The second DeepSeek shock: o1-class reasoning at a fraction of the price, weights open. Briefly wipes ~$600B off Nvidia's market cap — the moment markets priced in that inference was becoming cheap.
The largest single percentage cut on a flagship reasoning model — o3 drops to $2/$8 overnight, undercutting its own smaller sibling and confirming the reasoning tier follows the same deflation curve, just faster.
Llama-4-class, Qwen and DeepSeek models served by Together, Fireworks, Groq and DeepInfra put GPT-4-level capability below $1 per million tokens. Frontier closed models (GPT-5.5 at $5, Claude Opus tier) hold premium pricing — but the gap they must justify keeps widening.
What three years of data actually says
1. Capability tiers deflate ~10× per year; individual models rarely get cheaper. GPT-4o's list price was cut twice, but the bigger mechanism is replacement: the new model ships at the old model's lower tier. Budget for the tier, not the model name.
2. Output prices fall slower than input. GPT-4 launch output was 2× input ($60 vs $30); today 4–5× ratios are standard (GPT-4o: $10 vs $2.50). Providers protect margin on generation. If your workload is output-heavy — long answers, code generation — your effective deflation is meaningfully slower than headlines suggest.
3. Price rises happen. Claude 3.5 Haiku (+220%) is the famous one. When a "small" model benchmarks like a mid-tier, its price follows capability, not lineage. Never hardcode the assumption that the cheap tier stays cheap.
4. Shocks come from outside. Both discontinuous drops (May 2024, Jan 2025) came from DeepSeek, not from incumbent competition. The next repricing probably also arrives from somewhere unexpected — which argues for provider-agnostic plumbing over deep single-vendor integration.
What this means for your budget
Discount long-term projections. A feature that costs $10k/month in API calls today plausibly costs $1–3k/month at the same quality in 18 months. AI unit economics that look marginal now often work fine on the deflation curve — and committed annual contracts at today's rates fight that curve.
Re-shop quarterly. The 300+ model comparison table re-ranks every model by your actual token mix. Fifteen minutes a quarter routinely finds 30–60% savings from switching one workload tier down. The price-drop savings calculator prices exactly what a migration is worth.
Route by difficulty. The 100× spread between commodity and frontier pricing is the opportunity: most production queries don't need frontier. The model routing calculator shows the blended-cost math.
Live market check — right now, from OpenRouter
Fetching live model prices…
Live data from the public OpenRouter model list, fetched by your browser on page load — independent confirmation that the deflation curve on this page is still running.
🔔 Get an email when model prices change
We diff our 387-model price database on every refresh. When any provider moves a price — cuts or raises — subscribers get one plain-text email listing exactly what changed. No marketing, no schedule, only actual price events.