High volume tips toward self-hosting
Per-minute STT APIs are convenient; open models on a GPU win at scale. Estimate the GPU side on the GPU rental calculator before switching. For the reverse see the text-to-speech calculator.
What transcribing audio to text costs per month.
Per-minute STT APIs are convenient; open models on a GPU win at scale. Estimate the GPU side on the GPU rental calculator before switching. For the reverse see the text-to-speech calculator.
The Speech-to-Text API Cost Calculator estimates your monthly transcription bill from two inputs: the number of audio minutes transcribed per month and the price per minute charged by your provider. It multiplies these together to project a recurring cost, turning a per-minute rate that looks trivially small in isolation into a concrete monthly figure. The two main drivers are volume and unit price: because the total scales linearly with both, doubling either the minutes of audio you process or the rate you pay doubles the bill. This makes the tool useful for forecasting spend before you commit, comparing vendor rates, or sizing the impact of a growing user base.
The key trade-off to watch is that the per-minute rate often varies by model tier, language, and features, so the cheapest headline price may not reflect what you actually pay in production. A premium accuracy model or add-ons like speaker diarization or real-time streaming can carry a higher effective rate, while batch processing is frequently cheaper. A practical tip is to base your audio minutes per month on real usage rather than raw file duration, since silence, retries, and re-processing all count as billable time. Run the numbers at your expected peak volume, not just your average, so a growth spike does not produce a surprise invoice.
Per minute of audio processed, regardless of how many words it contains. Cost = minutes × per-minute rate. Add-ons like speaker diarisation, word timestamps and enhanced models can raise the rate.
For low volume, a hosted API is simplest and reliable. Above tens of thousands of minutes a month, running an open model such as Whisper on a rented GPU usually undercuts the per-minute price — the familiar build-vs-buy line.
Cost at different volumes:
| Volume | Monthly | Yearly |
|---|