Where the difference comes from
Input and output priced separately on both models — output usually drives the gap because it's the pricier side and workloads generate a lot of it.
| Side | Current | New | Change |
|---|
The bill change is the size of the prize
Switching models is one of the biggest levers on an AI bill, but the decision is usually made on a headline number — "this one is cheaper per million tokens" — that doesn't survive contact with a real workload. Your invoice is a blend of input and output pricing weighted by how your app actually uses the model, and output is where the money is: it's typically three to five times the input rate, and agents, chatbots and code generators produce a lot of it. Two models that look close on input can be twice as far apart once output is counted. This calculator prices both models on your real monthly input and output volumes and reports the difference as a dollar figure, a percentage, and a yearly number — the actual stake you're playing for.
Seeing that number changes the conversation. A migration that saves 8% a year rarely justifies the engineering time, the regression testing and the risk of a subtle quality drop. One that halves the bill is worth a serious evaluation. And the arrow points both ways: moving to a more capable, pricier model is easier to justify when you can see it only adds a manageable amount at your volume. Use this to size the prize first, then decide whether it's worth chasing — and pair it with the cheapest LLM API tool to scan the whole field and the reasoning token calculator if the new model thinks before it answers.
How to use it
1. Pick the model you're currently on and the one you're considering.
2. Enter your monthly input tokens (everything you send) and output tokens (everything the model writes), in millions.
3. Read the headline: green means the switch saves money, amber means it costs more.
4. Check the breakdown table to see whether input or output is driving the difference.
Common mistakes
Comparing on input price alone. Output is the pricier, higher-volume side for most apps — a model with cheap input and dear output can be the expensive choice. Assuming equal token counts. If the new model needs longer prompts or writes more to hit the same quality, its real cost is higher than this straight swap shows. Ignoring migration cost. Re-testing, prompt tuning and monitoring take engineering time; weigh it against the yearly saving. Forgetting discounts. Prompt caching and batch pricing can change the ranking — model both sides consistently.
FAQ
Where do I find my input and output token volumes?
Your provider's usage dashboard breaks the monthly total into input and output tokens. Divide each by a million and enter them here. If you only have a combined number, split it by your typical input-to-output ratio.
Does a cheaper model always cut my bill?
Only if it does the same job with the same tokens. A weaker model that needs more retries, longer prompts or more output can cost more in practice, so treat the saving here as the best case and validate quality before committing.
Can I compare within the same provider?
Yes — the two dropdowns are independent, so you can compare GPT-4o against GPT-4o mini, or a Pro tier against a Flash tier, exactly as you'd compare across providers.
Estimate only. Model prices change frequently — verify current per-million input and output rates before migrating.