8. Oktober 2026 · AI & LLMs · 4 Minuten Lesezeit
We wrote about Gemini 3.7 Flash's price doubling back in September — $0.75/$3.75 per million tokens through December 31, 2026, then $1.50/$7.50 from January 1, 2027. Since then, Google shipped 3.8 Flash, and it's reasonable to assume a newer model got a cleaner rate card. It didn't. I pulled the current Gemini API pricing docs directly and checked every Flash generation Google currently sells: 3.6 Flash, 3.7 Flash and 3.8 Flash all carry the exact same intro-rate clause, on the exact same date. If your plan was "move to the newest Flash before January and dodge the hike," that plan doesn't work — there's nowhere in the current Flash lineup left to move to.
| Modell | Jetzt, bis zum 31. Dezember 2026 | Ab 1. Januar 2027 | Ändern |
|---|---|---|---|
| Zwillinge 3.8 Blitz | $0.75 / $3.75 | $1.50 / $7.50 | 2.00x |
| Gemini 3.7 Flash | $0.75 / $3.75 | $1.50 / $7.50 | 2.00x |
| Zwillinge 3.6 Blitz | $0.75 / $3.75 | $1.50 / $7.50 | 2.00x |
| Gemini 3.5 Flash (keine geplante Änderung) | $1.50 / $9.00 | $1.50 / $9.00 | keine |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $0.30 / $2.50 | keine |
| Gemini 2.5 Flash (vorheriges Gen) | $0.30 / $2.50 | $0.30 / $2.50 | keine |
| Gemini 2.5 Flash-Lite | $0.10 / $0.40 | $0.10 / $0.40 | keine |
Three model names, three separate launch dates, one identical price line. That's not a coincidence of similar models landing on similar numbers — it's the same promotional structure applied to an entire model family: cheap intro rate to pull in adoption, doubling on the same January 1 date regardless of which specific Flash build you picked. Google's docs don't call out 3.6 or 3.8 by name the way the 3.7 announcement got attention — the clause just sits under the same rate card line for all three.
Eine Chatbot-App mit 300 MILLIONEN Eingabe-Token und 60 MILLIONEN Ausgabe-Token pro Monat — ein mittelgroßer Support-Bot, kein Edge-Case. Der Preis richtet sich nach dem aktuellen Flash-Modell, das gerade ausgeführt wird:
| Eingang (300 M) | Ausgang (60 M) | Monatliche Gesamtsumme | |
|---|---|---|---|
| Jetzt (3,6 / 3,7 / 3,8 Flash, bis 31. Dezember) | $225.00 | $225.00 | $450.00 |
| Ab 1. Januar 2027 (gleiches Modell, eines der drei) | $450.00 | $450.00 | $900.00 |
| Gemini 3.5 Flash-Lite, jederzeit (keine Wanderung) | $90.00 | $150.00 | $240.00 |
Der Sprung von $ 450 ist identisch, unabhängig davon, welche der drei aktuellen Flash-Generationen der Bot ausführt — innerhalb der Flash-Stufe selbst ist kein "Upgrade to dodge it" -Move verfügbar. Der einzige wirkliche Ausweg ist ein Stufenwechsel, kein Versionssprung: Flash-Lite führt heute, vor und nach Januar bereits genau diesen Workload für $ 240/Monat aus, weil er nie Teil der Aktionspreise war.
Don't spend engineering time migrating between 3.6, 3.7 and 3.8 Flash for cost reasons — they're the same bill, just with different model weights attached. If the January number matters to your budget, the decision that actually changes it is Flash vs. Flash-Lite (or staying on 2.5 Flash), decided on an eval set, not a model-version number. Run your real traffic through Flash-Lite before December 31 and measure whether answer quality holds — if it does, the $210/month difference in the example above is free money; if it doesn't, you now know to budget for the doubled rate with your eyes open instead of discovering it on the January invoice.
Neu bei der nutzungsbasierten LLM-Preisgestaltung? Beginnen Sie mit der kostenlose API-Kostenleitfäden.
Stellen Sie es selbst bereit: DigitalOcean – 200 $ kostenloses Guthaben ↗ - Hostinger VPS ↗
Die Preise wurden mit den offiziellen Preisdokumenten der Gemini-API von Google (ai.google.dev/gemini-api/docs/pricing) verglichen und am 08.10.2026 überprüft. Referenzschätzungen unter Verwendung der von Google veröffentlichten Pay-as-you-go-Raten — Vertex AI-Preise, Unternehmensvereinbarungen und zukünftige Preisänderungen variieren; bestätigen Sie die aktuelle Preisgestaltung an der Quelle vor der Budgetierung. Workload-Zahlen (Token-Zählungen) sind veranschaulichende Szenarien, die anhand realer veröffentlichter pro-Token-Raten berechnet werden, nicht anhand der vom Lieferanten bereitgestellten Zahlen.