Home › Blog › Gemini Flash Preiserhöhung trifft drei Modelle

Das Wechseln der Gemini-Blitzversionen wird der Preiserhöhung im Januar nicht ausweichen

8. Oktober 2026 · AI & LLMs · 4 Minuten Lesezeit

We wrote about Gemini 3.7 Flash's price doubling back in September — $0.75/$3.75 per million tokens through December 31, 2026, then $1.50/$7.50 from January 1, 2027. Since then, Google shipped 3.8 Flash, and it's reasonable to assume a newer model got a cleaner rate card. It didn't. I pulled the current Gemini API pricing docs directly and checked every Flash generation Google currently sells: 3.6 Flash, 3.7 Flash and 3.8 Flash all carry the exact same intro-rate clause, on the exact same date. If your plan was "move to the newest Flash before January and dodge the hike," that plan doesn't work — there's nowhere in the current Flash lineup left to move to.

Jedes aktuelle Flash-Modell, Seite an Seite

ModellJetzt, bis zum 31. Dezember 2026Ab 1. Januar 2027Ändern
Zwillinge 3.8 Blitz$0.75 / $3.75$1.50 / $7.502.00x
Gemini 3.7 Flash$0.75 / $3.75$1.50 / $7.502.00x
Zwillinge 3.6 Blitz$0.75 / $3.75$1.50 / $7.502.00x
Gemini 3.5 Flash (keine geplante Änderung)$1.50 / $9.00$1.50 / $9.00keine
Gemini 3.5 Flash-Lite$0.30 / $2.50$0.30 / $2.50keine
Gemini 2.5 Flash (vorheriges Gen)$0.30 / $2.50$0.30 / $2.50keine
Gemini 2.5 Flash-Lite$0.10 / $0.40$0.10 / $0.40keine

Three model names, three separate launch dates, one identical price line. That's not a coincidence of similar models landing on similar numbers — it's the same promotional structure applied to an entire model family: cheap intro rate to pull in adoption, doubling on the same January 1 date regardless of which specific Flash build you picked. Google's docs don't call out 3.6 or 3.8 by name the way the 3.7 announcement got attention — the clause just sits under the same rate card line for all three.

Was das bei einer echten Arbeitsbelastung tatsächlich kostet

Eine Chatbot-App mit 300 MILLIONEN Eingabe-Token und 60 MILLIONEN Ausgabe-Token pro Monat — ein mittelgroßer Support-Bot, kein Edge-Case. Der Preis richtet sich nach dem aktuellen Flash-Modell, das gerade ausgeführt wird:

Eingang (300 M)Ausgang (60 M)Monatliche Gesamtsumme
Jetzt (3,6 / 3,7 / 3,8 Flash, bis 31. Dezember)$225.00$225.00$450.00
Ab 1. Januar 2027 (gleiches Modell, eines der drei)$450.00$450.00$900.00
Gemini 3.5 Flash-Lite, jederzeit (keine Wanderung)$90.00$150.00$240.00

Der Sprung von $ 450 ist identisch, unabhängig davon, welche der drei aktuellen Flash-Generationen der Bot ausführt — innerhalb der Flash-Stufe selbst ist kein "Upgrade to dodge it" -Move verfügbar. Der einzige wirkliche Ausweg ist ein Stufenwechsel, kein Versionssprung: Flash-Lite führt heute, vor und nach Januar bereits genau diesen Workload für $ 240/Monat aus, weil er nie Teil der Aktionspreise war.

Was ich eigentlich tun würde

Don't spend engineering time migrating between 3.6, 3.7 and 3.8 Flash for cost reasons — they're the same bill, just with different model weights attached. If the January number matters to your budget, the decision that actually changes it is Flash vs. Flash-Lite (or staying on 2.5 Flash), decided on an eval set, not a model-version number. Run your real traffic through Flash-Lite before December 31 and measure whether answer quality holds — if it does, the $210/month difference in the example above is free money; if it doesn't, you now know to budget for the doubled rate with your eyes open instead of discovering it on the January invoice.

Neu bei der nutzungsbasierten LLM-Preisgestaltung? Beginnen Sie mit der kostenlose API-Kostenleitfäden.

Stellen Sie es selbst bereit: DigitalOcean – 200 $ kostenloses Guthaben ↗ - Hostinger VPS ↗

Die Preise wurden mit den offiziellen Preisdokumenten der Gemini-API von Google (ai.google.dev/gemini-api/docs/pricing) verglichen und am 08.10.2026 überprüft. Referenzschätzungen unter Verwendung der von Google veröffentlichten Pay-as-you-go-Raten — Vertex AI-Preise, Unternehmensvereinbarungen und zukünftige Preisänderungen variieren; bestätigen Sie die aktuelle Preisgestaltung an der Quelle vor der Budgetierung. Workload-Zahlen (Token-Zählungen) sind veranschaulichende Szenarien, die anhand realer veröffentlichter pro-Token-Raten berechnet werden, nicht anhand der vom Lieferanten bereitgestellten Zahlen.