Tip: a chat product typically runs 3–6× more input than output tokens (history is re-sent each turn). If you only know total tokens, split them 80/20.

0 models fit · sorted by quality tier, then headroom ·

ModelProviderMonthly cost% of budgetHeadroomTier

Prices are reference estimates (). Report outdated price →

How to use the headroom column

Headroom is how many times your current volume could grow before the model blows the budget. A model at 45% of budget has ~2.2× headroom; at 90%, only 1.1× — the first growth month breaks it. LLM usage almost always grows faster than planned (longer conversations, retries, new features quietly added), so treat 2× headroom as the practical minimum for a production pick.

The smart play is usually a pair: a cheap workhorse inside budget with 5×+ headroom for the 80% of easy queries, plus a flagship used sparingly for the hard 20%. The model routing calculator shows the blended math, and the full comparison table lets you inspect any model this finder surfaces. Prices here fall ~10×/year at constant capability — see the price history tracker — so re-run this check quarterly; the set of models that fit your budget grows on its own.