blended monthly cost
all-standard cost
saved / mo
blended $/1M

Where the money goes — per tier

Your traffic split across the three tiers. Whatever is not deferrable or critical stays on standard. Batch/flex is charged at the discount; priority carries the premium.

TierShareRate vs standardMonthly cost

How much can you batch? — the deferrable dial

The single biggest lever is the share of volume you can move off standard onto the discounted tier. Each row holds your priority share fixed and varies the deferrable share; the saving is versus running everything on standard.

Deferrable shareBlended costSaved / moSaving

The cheapest token is the one you were willing to wait for

Most teams treat model price as a single number, but the same tokens now come in tiers, and the gap between them is enormous: roughly half price for anything you can defer, a premium for anything you insist on serving instantly. The mistake in both directions is uniformity — running everything on standard leaves the batch discount on the table, and running everything on priority pays a latency tax on calls no user is waiting for. The win is almost always in the boring middle: tag your traffic by how much delay it can absorb, push the offline half onto batch or flex, keep ordinary interactive work on standard, and reserve priority for the narrow set of paths where a slow response actually costs money. The priority premium is the number people misjudge most, so this calculator isolates it — it prints the extra dollars per month you are paying purely for the service level on your critical share, separate from the tokens themselves, so the decision is a comparison and not a shrug. Size the offline discount precisely on the batch API savings calculator, stack it with cache hits on the prompt caching savings calculator, and fold every lever together on the LLM cost optimization calculator.

Host your project:DigitalOcean — $200 free ↗Hostinger VPS
Batch API SavingsPrompt Caching SavingsLLM Cost OptimizationLLM Tier PricingLLM Latency