8 de octubre de 2026 · AI y LLM · 4 min de lectura
We wrote about Gemini 3.7 Flash's price doubling back in September — $0.75/$3.75 per million tokens through December 31, 2026, then $1.50/$7.50 from January 1, 2027. Since then, Google shipped 3.8 Flash, and it's reasonable to assume a newer model got a cleaner rate card. It didn't. I pulled the current Gemini API pricing docs directly and checked every Flash generation Google currently sells: 3.6 Flash, 3.7 Flash and 3.8 Flash all carry the exact same intro-rate clause, on the exact same date. If your plan was "move to the newest Flash before January and dodge the hike," that plan doesn't work — there's nowhere in the current Flash lineup left to move to.
| Modelo | Ahora, hasta el 31 de diciembre de 2026 | A partir del 1 de enero de 2027 | Cambiar |
|---|---|---|---|
| Flash Géminis 3.8 | $0.75 / $3.75 | $1.50 / $7.50 | 2.00x |
| Flash de Géminis 3.7 | $0.75 / $3.75 | $1.50 / $7.50 | 2.00x |
| Flash Géminis 3.6 | $0.75 / $3.75 | $1.50 / $7.50 | 2.00x |
| Flash Géminis 3.5 (sin cambio programado) | $1.50 / $9.00 | $1.50 / $9.00 | nada |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | $0.30 / $2.50 | nada |
| Géminis 2.5 Flash (generación anterior) | $0.30 / $2.50 | $0.30 / $2.50 | nada |
| Gemini 2.5 Flash-Lite | $0.10 / $0.40 | $0.10 / $0.40 | nada |
Three model names, three separate launch dates, one identical price line. That's not a coincidence of similar models landing on similar numbers — it's the same promotional structure applied to an entire model family: cheap intro rate to pull in adoption, doubling on the same January 1 date regardless of which specific Flash build you picked. Google's docs don't call out 3.6 or 3.8 by name the way the 3.7 announcement got attention — the clause just sits under the same rate card line for all three.
Una aplicación de chatbot que ejecuta 300 MILLONES de tokens de entrada y 60 MILLONES de tokens de salida al mes: un bot de soporte de tamaño mediano, no un caso extremo. Con un precio en cualquier modelo Flash de generación actual que se esté ejecutando:
| Entrada (300M) | Salida (60M) | Total mensual | |
|---|---|---|---|
| Ahora (Flash 3.6 / 3.7 / 3.8, hasta el 31 de diciembre) | $225.00 | $225.00 | $450.00 |
| A partir del 1 de enero de 2027 (mismo modelo, cualquiera de los tres) | $450.00 | $450.00 | $900.00 |
| Gemini 3.5 Flash-Lite, en cualquier momento (sin caminata) | $90.00 | $150.00 | $240.00 |
El salto de $ 450 es idéntico sin importar cuál de las tres generaciones actuales de Flash esté ejecutando el bot: no hay un movimiento de "actualización para esquivarlo" disponible dentro del propio nivel de Flash. La única salida real es un cambio de nivel, no un aumento de versión: Flash-Lite ya ejecuta esta carga de trabajo exacta por $ 240/mes hoy, antes y después de enero, porque nunca fue parte de los precios promocionales en primer lugar.
Don't spend engineering time migrating between 3.6, 3.7 and 3.8 Flash for cost reasons — they're the same bill, just with different model weights attached. If the January number matters to your budget, the decision that actually changes it is Flash vs. Flash-Lite (or staying on 2.5 Flash), decided on an eval set, not a model-version number. Run your real traffic through Flash-Lite before December 31 and measure whether answer quality holds — if it does, the $210/month difference in the example above is free money; if it doesn't, you now know to budget for the doubled rate with your eyes open instead of discovering it on the January invoice.
¿Nuevo en los precios de LLM basados en el uso? Comience con el guías gratuitas de costos de API.
Despliéguelo usted mismo: DigitalOcean — $200 de crédito gratis ↗ · VPS Hostinger ↗
Los precios se verificaron con los documentos oficiales de precios de la API de Gemini de Google (ai.google.dev/gemini-api/docs/pricing), verificados el 08-10-2026. Estimaciones de referencia utilizando las tarifas de pago por uso publicadas de Google: los precios de Vertex AI, los acuerdos empresariales y los cambios futuros de tarifas varían; confirme los precios actuales en la fuente antes de presupuestar. Las cifras de carga de trabajo (recuentos de tokens) son escenarios ilustrativos calculados contra las tasas reales publicadas por token, no los números suministrados por el proveedor.