Published 15 July 2026 · reference prices, verify before budgeting
"How much will our AI support chatbot cost?" is one of the most common questions we see. The answer varies by 23× depending on the model — and grows faster than you expect as conversation history accumulates.
| Model | $/1M in | $/1M out | Cost per chat | 100k chats/mo | Annual |
|---|---|---|---|---|---|
| Gemini 2.0 Flash | $0.10 | $0.40 | $0.0043 | $432 | $5,184 |
| GPT-4o mini | $0.15 | $0.60 | $0.0065 | $648 | $7,776 |
| Claude Haiku 4.5 | $0.80 | $4.00 | $0.0374 | $3,744 | $44,928 |
| GPT-4o | $2.50 | $10.00 | $0.1013 | $10,125 | $121,500 |
| Claude Sonnet 4 | $3.00 | $15.00 | $0.1258 | $12,576 | $150,912 |
This is where most chatbot cost estimates go wrong. They model cost per message at a fixed token count — but in reality, each message carries the full history.
| Message # | Input tokens (incl. history) | Output tokens | Flash cost | GPT-4o cost |
|---|---|---|---|---|
| 1 | 550 | 280 | $0.000167 | $0.00418 |
| 2 | 900 | 280 | $0.000202 | $0.00505 |
| 4 | 1,400 | 280 | $0.000252 | $0.00630 |
| 6 | 1,950 | 280 | $0.000307 | $0.00768 |
| 8 (final) | 2,500 | 280 | $0.000362 | $0.00905 |
| Total (8 msg) | ~10,800 | ~2,240 | $0.0043 | $0.101 |
Not all support queries need the same model. Route queries to different tiers:
A simple intent classifier (itself cheap — Flash at ~$0.002 per routing call) can gate which tier handles each conversation.
Instead of including the full transcript in every message, summarize the conversation after message 4 and pass a compressed summary (≈200 tokens) for messages 5+. This keeps context stable rather than growing linearly.
Effect: reduces average input per message from 1,350 to ~600 tokens — a 55% cut in input costs. On GPT-4o: $10,125 → ~$5,400/month.
Your system prompt and FAQ context (400 tokens in our model) is repeated in every single message. Both Anthropic and OpenAI offer prompt caching — repeated prefix tokens are billed at 10–50% of the normal rate after the first request.
If your system prompt is 1,000 tokens and you cache it, you save 90% on those 1,000 tokens across all subsequent messages. On 100k chats × 8 messages = 800k messages, that's 800M tokens × $0.10/1M × 90% = $72/month saved on Flash alone. On GPT-4o the saving is larger in absolute terms.
| Scale | Flash | Blended (routing) | GPT-4o |
|---|---|---|---|
| 10k chats/mo (startup) | $43 | $164 | $1,013 |
| 100k chats/mo (growth) | $432 | $1,639 | $10,125 |
| 500k chats/mo (scale) | $2,160 | $8,195 | $50,625 |
| 1M chats/mo (enterprise) | $4,320 | $16,390 | $101,250 |
At 1M chats/month, the difference between Flash and GPT-4o is $97k/month — $1.16M per year. Even if Flash requires 15% more escalations to human agents, the math almost always favours the cheaper model tier.
That's the real question, and it depends entirely on your use case. For FAQ answering, order status, return policy, shipping info — Flash performs comparably to GPT-4o in controlled evaluations. For nuanced complaint handling, multi-step troubleshooting, or situations requiring genuine reasoning about edge cases — frontier models have a measurable advantage.
The practical approach: run Flash for 30 days, measure escalation rate and CSAT, then compare to your baseline. If CSAT drops <1 point and escalation rises <5%, the model switch is justified by the cost saving.
At 100,000 chats/month with 8 messages average: Gemini Flash ~$432, GPT-4o mini ~$648, Claude Haiku ~$3,744, GPT-4o ~$10,125. The model choice is the biggest cost lever.
Every message includes the full history as input tokens. Message 1 costs ~400 input tokens; message 8 costs ~2,500 — a 6× increase for the same output. Without conversation summarization, costs grow linearly with chat length.
For FAQ answering, order status, and policy questions — typically yes. For nuanced complaint handling or complex troubleshooting — benchmark first. Many teams run Flash for 60-70% of volume and escalate to a frontier model for edge cases.
Three main levers: (1) model routing — use Flash for simple queries, GPT-4o only for complex ones; (2) conversation summarization — compress old history into ~200 tokens; (3) prompt caching — both Anthropic and OpenAI offer 50-90% discounts on repeated prompt prefixes.