Published 2026-07-11 · based on our new LLM price history tracker — every major price event since GPT-4, with dates
We compiled every major LLM API price change from March 2023 to today into one timeline. The pattern is more consistent than we expected, and it breaks most of the AI budget spreadsheets we've seen: at constant capability, prices fall roughly 10× per year. Not gradually — in steps, and usually from directions nobody predicted.
| Date | Frontier-class input, $/1M | vs GPT-4 launch |
|---|---|---|
| Mar 2023 (GPT-4) | $30.00 | — |
| Aug 2024 (GPT-4o after cut) | $2.50 | −92% |
| 2026 (open-weight hosts) | <$1.00 | −97% |
Historical list prices. Full timeline with 14 dated events on the price history page.
The mechanism matters: individual models rarely get cheaper (GPT-4o's two cuts were the exception). What actually happens is replacement — each new generation ships at the previous generation's lower tier. GPT-4o mini beat GPT-3.5 Turbo at a third of its price. DeepSeek V2 delivered near-frontier quality at $0.14. You don't ride the deflation by waiting for discounts; you ride it by re-shopping the tier.
Annual contracts at today's rates. A 30% volume discount locked for 12 months is a loss if market prices fall 60–90% in that window — and historically, they have. Anything longer than a 6–12 month commit is a bet against a three-year trend. If you must commit, commit to spend, not to a specific model's rate card.
"AI is too expensive for this feature" decisions. A feature that pencils out at $10k/month today plausibly costs $1–3k at equal quality inside 18 months. Marginal unit economics now usually means good unit economics soon — the build-vs-wait calculus should include the curve, not just today's sticker.
Deep single-vendor integration. Both discontinuous price shocks in our timeline (May 2024, Jan 2025) came from DeepSeek — an outsider. Teams with provider-agnostic routing captured those prices in days; teams welded to one SDK watched from the sidelines. The engineering cost of an abstraction layer is an insurance premium on the next shock.
Two honest caveats from the data. First, output prices fall slower — GPT-4 launched with output at 2× input; today 4–5× ratios are standard. Output-heavy workloads (code generation, long answers) deflate meaningfully slower than the headline. Second, prices do rise: Claude 3.5 Haiku launched at 3.2× its predecessor's price because it benchmarked like a mid-tier model. The cheap tier staying cheap is a habit, not a guarantee — never hardcode it.
Quarterly, fifteen minutes: put your real token mix into the 300+ model comparison and check whether your current model still wins its tier. When it doesn't, the price-drop savings calculator prices the migration. Most teams we model find 30–60% savings the first time they do this — money that was sitting on the table simply because the model choice was made two quarters ago.
All figures are historical list prices, compiled July 2026 — details and sources on the price history page. Related: AI Model Price Comparison · Model Routing Savings · Cheapest LLM API.