real blended monthly cost
naive flat-rate estimate
hidden surcharge exposure
savings if long requests trimmed
Surcharge exposure

Cost breakdown at your workload

Short requests price normally at the base rate; long requests price their ENTIRE call — input and output — at the higher cliff rate. The naive row shows what a flat-rate estimate would assume, and the trimmed row shows the scenario where long requests are compressed to just under the threshold.

RowRequestsCost/requestSubtotal

Blended cost by share of long requests

Everything else held at the values above, swept across the share of requests that land over the threshold. The highlighted row is closest to your current setting.

Over-threshold shareActual costNaive estimateSurcharge

How this connects to other tools

This calculator prices one specific mechanic: a hard length threshold where crossing it reprices an entire request, not just the tokens over the line. That's different from our context window cost calculator, which models a smooth, flat per-token cost by context size and has no cliff or threshold logic at all — use that tool if the model you're pricing charges the same rate regardless of request length. It's also different from the LLM tier pricing calculator and LLM service tier calculator, which price latency-based Batch/Flex/Priority service tiers — a choice you make per request about how fast you need a response, not a consequence of how long your prompt is. If you want general Gemini pricing across model sizes without the cliff mechanic isolated out, see the Gemini API cost calculator. This page exists because a flat per-token view of Gemini 2.5 Pro's headline $1.25 rate misses the real cost entirely once any share of a workload crosses 200K input tokens.

Host your project:DigitalOcean — $200 free ↗Hostinger VPS