How this calculator works
Free LLM service tier calculator โ your tokens and standard price to the cost on Batch, Flex, Standard and Priority tiers.
Frequently asked questions
What are LLM service tiers?
Most big providers sell the same model at several latency tiers. Batch runs your requests asynchronously (often within 24h) for about half price; a Flex or service tier is cheaper but slower or variable; Standard is the default; a Priority tier costs a premium for guaranteed fast responses.
When should I use the Batch tier?
Whenever the work is not user-facing in real time โ bulk summarisation, embeddings, evals, data labelling, overnight jobs. Batch is typically ~50% off, so moving async workloads off the standard tier can roughly halve that part of the bill for zero quality loss.