HomeBlog › LLM price deflation and your budget

Your AI budget is probably wrong — LLM prices fall ~10× a year

Published 2026-07-11 · based on our new LLM price history tracker — every major price event since GPT-4, with dates

We compiled every major LLM API price change from March 2023 to today into one timeline. The pattern is more consistent than we expected, and it breaks most of the AI budget spreadsheets we've seen: at constant capability, prices fall roughly 10× per year. Not gradually — in steps, and usually from directions nobody predicted.

Three numbers that tell the story

DateFrontier-class input, $/1Mvs GPT-4 launch
Mar 2023 (GPT-4)$30.00
Aug 2024 (GPT-4o after cut)$2.50−92%
2026 (open-weight hosts)<$1.00−97%

Historical list prices. Full timeline with 14 dated events on the price history page.

The mechanism matters: individual models rarely get cheaper (GPT-4o's two cuts were the exception). What actually happens is replacement — each new generation ships at the previous generation's lower tier. GPT-4o mini beat GPT-3.5 Turbo at a third of its price. DeepSeek V2 delivered near-frontier quality at $0.14. You don't ride the deflation by waiting for discounts; you ride it by re-shopping the tier.

What this breaks

Annual contracts at today's rates. A 30% volume discount locked for 12 months is a loss if market prices fall 60–90% in that window — and historically, they have. Anything longer than a 6–12 month commit is a bet against a three-year trend. If you must commit, commit to spend, not to a specific model's rate card.

"AI is too expensive for this feature" decisions. A feature that pencils out at $10k/month today plausibly costs $1–3k at equal quality inside 18 months. Marginal unit economics now usually means good unit economics soon — the build-vs-wait calculus should include the curve, not just today's sticker.

Deep single-vendor integration. Both discontinuous price shocks in our timeline (May 2024, Jan 2025) came from DeepSeek — an outsider. Teams with provider-agnostic routing captured those prices in days; teams welded to one SDK watched from the sidelines. The engineering cost of an abstraction layer is an insurance premium on the next shock.

What the trend does NOT say

Two honest caveats from the data. First, output prices fall slower — GPT-4 launched with output at 2× input; today 4–5× ratios are standard. Output-heavy workloads (code generation, long answers) deflate meaningfully slower than the headline. Second, prices do rise: Claude 3.5 Haiku launched at 3.2× its predecessor's price because it benchmarked like a mid-tier model. The cheap tier staying cheap is a habit, not a guarantee — never hardcode it.

The practical routine

Quarterly, fifteen minutes: put your real token mix into the 300+ model comparison and check whether your current model still wins its tier. When it doesn't, the price-drop savings calculator prices the migration. Most teams we model find 30–60% savings the first time they do this — money that was sitting on the table simply because the model choice was made two quarters ago.

All figures are historical list prices, compiled July 2026 — details and sources on the price history page. Related: AI Model Price Comparison · Model Routing Savings · Cheapest LLM API.