reserved / month
vs on-demand
effective $/1M
capacity used

Effective price as volume grows

The reservation cost is fixed, so every extra token drives the effective per-million price down — until you hit the capacity ceiling and overflow bills at the on-demand rate.

Monthly volumeOn-demand costReserved effective $/1MCheaper

Reserve capacity only above the break-even

Provisioned throughput — PTUs on Azure OpenAI, provisioned throughput on AWS Bedrock — swaps a per-token bill for a fixed monthly reservation of dedicated capacity. It buys predictable latency and a flat cost, but it only saves money once your steady token volume clears the break-even: the reservation cost divided by the on-demand price per token. Below that you are paying for idle units; above it the fixed cost spreads over more tokens and your effective price per million keeps falling, right up to the capacity ceiling of units × tokens-per-minute × minutes in the month. Push past the ceiling and the overflow bills at full on-demand rate, quietly eroding the discount. Size the reservation to the volume you are confident you will use every month, forecast growth on the API cost forecast calculator, and compare a spend commitment on the committed spend discount calculator.

Host your project:DigitalOcean — $200 free ↗Hostinger VPS
LLM Evaluation Cost CalculatorSynthetic Data Generation Cost CalculatorAI Seat License vs API Cost CalculatorBatch Processing Time & Cost CalculatorAI Cost Per User Calculator