Your effective rate depends on your token mix
Split input/output pricing means the sticker rate is not your rate โ output usually dominates. Cap max_tokens to control it. Turn the per-call number into a full bill on the LLM token cost calculator.
Input and output are priced differently โ get your true effective rate.
Split input/output pricing means the sticker rate is not your rate โ output usually dominates. Cap max_tokens to control it. Turn the per-call number into a full bill on the LLM token cost calculator.
Free blended LLM price calculator โ input and output prices and token mix to your real per-call cost and effective per-1M rate.
Generating tokens is more compute-intensive than reading them, so providers price output roughly 3โ5ร higher than input. A prompt with a huge context but a short answer can still be dominated by the few expensive output tokens โ or the reverse, depending on your token mix.
Your true effective rate once input and output are weighted by how many of each you actually use. Advertised prices are split; the blended rate tells you what a real call costs per token, which is what you should use for margin and forecasting.