18 September 2026 · AI & LLMs · 6 min read
When Claude Sonnet 5 launched, Anthropic billed $2 per million input tokens and $10 per million output as introductory pricing — through August 31, 2026, with a step up to $3/$15 scheduled for September 1. I went looking for that increase to update our own pricing page, because it should have landed over two weeks ago. It didn't. Anthropic's official pricing docs now carry a note saying the hike "will not occur" and the $2/$10 launch price is permanent. So I pulled the full Claude 5 family rate card — Fable 5, Opus 5, Sonnet 5, Haiku 4.5 — straight from the source and ran real workloads through it to see what that cancellation is actually worth, and how the four tiers compare once you're not just reading a table.
Take a mid-size product: a support copilot doing 2,000 requests a day, each carrying roughly 8,000 tokens of retrieved context and conversation history in, and returning a 500-token answer. That's 480 million input tokens and 30 million output tokens a month — not a huge company, just a real feature with real traffic.
| Sonnet 5 rate | Input (480M) | Output (30M) | Monatliche Gesamtsumme |
|---|---|---|---|
| Actual: $2 / $10 per MTok | $960.00 | $300.00 | $1,260.00 |
| Cancelled: $3 / $15 per MTok | $1,440.00 | $450.00 | $1,890.00 |
That's $630 a month, $7,560 a year, that this product isn't paying because Anthropic backed off the increase. Nobody sent an invoice for it, nobody wrote a change request — the bill just stayed where it was. If you budgeted for the September step-up like the original announcement told you to, you can put that line back.
Verified against Anthropic's own pricing docs, checked 2026-09-18. This is the complete current lineup plus the two previous-gen models still served for existing integrations.
| Modell | Eingang | 5m cache write | 1h cache write | Cache-Treffer | Ausgabe |
|---|---|---|---|---|---|
| Claude Fabel 5 | $10.00 | $12.50 | $20.00 | $1.00 | $50.00 |
| Claude Opus 5 | $5.00 | $6.25 | $10.00 | $0.50 | $25.00 |
| Claude Opus 4.8 (gleicher Satz) | $5.00 | $6.25 | $10.00 | $0.50 | $25.00 |
| Claude Sonett 5 | $2.00 | $2.50 | $4.00 | $0.20 | $10.00 |
| Claude Sonett 4.6 (prior gen) | $3.00 | $3.75 | $6.00 | $0.30 | $15.00 |
| Claude Haiku 4.5 | $1.00 | $1.25 | $2.00 | $0.10 | $5.00 |
Alle Preise pro Million Token. Bemerkenswert: Sonnet 5 bei $ 2/$ 10 unterbietet jetzt seine eigene vorherige Generation, Sonnet 4.6 bei $ 3/$ 15 — das neue Modell ist sowohl besser als auch 33% billiger als das, das es ersetzt hat, was fast nie bei einer Markteinführung der Fall ist.
Ich habe die gleiche monatliche Arbeitslast von 480 M Input/30 M Output durch alle vier aktuellen Modelle geleitet, ohne weitere Änderungen. Der Aufstrich überraschte mich weniger für seine Größe als dafür, wie aufgeräumt er ist.
| Modell | Monatliche Kosten | vs Haiku 4.5 |
|---|---|---|
| Claude Haiku 4.5 | $630.00 | 1.0x |
| Claude Sonett 5 | $1,260.00 | 2.0x |
| Claude Opus 5 | $3,150.00 | 5.0x |
| Claude Fabel 5 | $6,300.00 | 10.0x |
Sonnet costs exactly double Haiku, Opus exactly five times Haiku, Fable exactly ten. Because input and output are both round multiples of Haiku's rate across every tier, the ladder holds for any workload shape, not just this one — which makes back-of-envelope tier comparisons genuinely reliable instead of a rough guess. The real decision isn't "which is cheapest," it's whether the task in front of you needs Opus-or-Fable-grade reasoning, because the cost difference for getting that wrong compounds fast at volume.
Angenommen, Sie führen ein Coding-Agent-Produkt aus, das den gleichen 6.000-Token-Repo-Kontext bei jeder Folgeumdrehung innerhalb einer Sitzung erneut sendet, und eine Sitzung dreht durchschnittlich 20 Umdrehungen innerhalb eines 5-Minuten-Fensters — genau das, wofür der 5-Minuten-Cache von Anthropic gebaut ist. Auf Sonnet 5: Ein 5-minütiger Cache-Schreibvorgang kostet 2,50 $/MTok (1,25x Basiseingabe), ein Cache-Treffer kostet 0,20 $/MTok (0,1 x Basiseingabe).
| Pro 20-Turn-Session | Bei 50.000 Sitzungen/Monat | |
|---|---|---|
| Kein Caching (120.000 ctx-Token × $ 2/MTok-Basis) | $0.2400 | $12,000.00 |
| Mit Zwischenspeicherung (1 Schreiben + 19 Treffer) | $0.0378 | $1,890.00 |
Das sind 10.110 $ pro Monat, eine 84% ige Kürzung, weil nichts am Produkt geändert wurde — nur das Caching für den Kontext einzuschalten, der bereits wortwörtlich zurückgeschickt wurde. Wenn Ihr Agent oder Ihre RAG-App DEN wiederholten Kontext noch nicht zwischenspeichert, ist dies der einzelne Posten mit der höchsten Hebelwirkung auf der gesamten Preiskarte.
Für alles, was keine Live-Antwort erfordert — Klassifizierung über Nacht, Massen-Tagging, Nachfüllungen — bietet die Batch-API einen pauschalen Rabatt von 50 % auf Input und Output, auf jedes Modell, ohne Ausnahmen. Ein 100M-Input /100M-Output-Klassifizierungslauf auf Haiku 4.5 kostet 600 $ zu Standardtarifen und genau 300 $ gebündelt. Es ist die einfachste Einsparung auf dieser ganzen Seite: keine Architekturänderung, nur ein anderer Endpunkt und ein asynchrones Ergebnis.
Default to Sonnet 5 unless you've specifically measured that a task needs more. At $2/$10 it now undercuts the model it replaced while scoring higher on Anthropic's own benchmarks, which is the rare case where "use the new default" is also the cheap option. Reach for Opus 5 only where the reasoning gap is visible in your own evals — not because it feels safer — and reserve Fable 5 for the genuinely hard, long-horizon runs where a 10x bill over Haiku is obviously worth it. And turn prompt caching on before you optimize anything else; in the example above it beat every model-tier decision on this page.
Neu bei der nutzungsbasierten LLM-Preisgestaltung? Beginnen Sie mit der kostenlose API-Kostenleitfäden.
Stellen Sie es selbst bereit: DigitalOcean – 200 $ kostenloses Guthaben ↗ - Hostinger VPS ↗
Pricing verified live against Anthropic's official Claude API pricing docs (platform.claude.com/docs/en/about-claude/pricing), checked 2026-09-18. Reference estimates using Anthropic's published self-serve rates — Bedrock, Google Cloud and Microsoft Foundry pricing, enterprise volume discounts and future rate changes vary; confirm current pricing at the source before budgeting. Workload figures (request volumes, token counts) are illustrative scenarios computed against real published per-token rates, not vendor-supplied numbers.