Home โ€บ Blog โ€บ Speech-to-text cost per hour

What Speech-to-Text Actually Costs Per Hour of Audio (2026)

26 July 2026 ยท AI & LLMs ยท 5 min read

I priced out transcription for a side project last week and the thing that surprised me was how small the numbers are. A whole hour of audio turned into text costs somewhere between a quarter and forty cents on the mainstream APIs. Not per minute. Per hour. The expensive part of a voice pipeline is almost never the transcription โ€” it is everything you bolt on around it. But the per-hour rates still matter once you are running hundreds of hours a month, so here is the real math.

The real 2026 prices

These are the pay-as-you-go rates I track for the three APIs most people actually reach for. I converted everything to a per-hour figure so you can compare apples to apples โ€” some vendors quote per minute, some per hour, and that alone hides the difference.

APIQuoted ratePer hour of audioFree credit
Deepgram (Nova)$0.0043 / min~$0.26$200 signup credit
OpenAI Whisper$0.006 / min$0.36โ€”
AssemblyAI (Universal)$0.37 / hr$0.37Free hours to start

So Deepgram's batch model is the cheapest of the three at roughly $0.26 an hour, Whisper sits in the middle at $0.36, and AssemblyAI's Universal model is $0.37. Streaming (live) transcription costs a little more than batch on every provider, so if you need real-time captions, budget above these numbers.

What a real workload costs

Rates per hour feel abstract, so here is what three actual volumes cost per month. Say you are transcribing support calls or a podcast back-catalogue:

Audio / monthDeepgramWhisperAssemblyAI
100 hours$25.80$36.00$37.00
500 hours$129.00$180.00$185.00
2,000 hours$516.00$720.00$740.00

At 500 hours a month the gap between cheapest and priciest is about $56. Real, but not the thing that will sink your budget. At 2,000 hours it grows to roughly $224/month. That is the point where picking Deepgram over the others starts to pay for the migration effort โ€” and not really before.

Where the money actually hides

Here is what I learned the expensive way. The transcription line item is small; the surrounding features are where the bill balloons. Speaker labels (diarization), sentiment, summarization, PII redaction and language detection are usually priced as add-ons, and each one stacks on top of the base rate. Turn on three of them and your $0.37/hr quietly becomes $0.70+/hr. AssemblyAI in particular bundles a lot of intelligence models that are billed separately once you enable them.

The other trap is that you get charged for audio duration, not word count. Silence, hold music and cross-talk all bill at the full rate. A one-hour recording with 20 minutes of dead air still costs you a full hour. Trimming silence before you upload is the single cheapest optimization nobody does.

What I actually default to

For batch jobs where I control the audio, Deepgram Nova wins on price and it is fast. When I am already inside an OpenAI stack and volume is modest, Whisper at $0.006/min is not worth switching away from โ€” the $200-ish annual difference is smaller than a day of engineering time. AssemblyAI earns its slightly higher rate only if you genuinely want its built-in summarization and speaker analytics instead of wiring your own. Run your own hours through the numbers before you commit:

New to usage-based API pricing? Start with the free API-cost guides. And if you are building a full voice bot, transcription is only one third of the bill โ€” see what 100k support chats actually cost for the LLM side.

Prices are provider list rates as tracked in our catalog (July 2026) and are reference estimates โ€” confirm current rates on each provider's pricing page before budgeting. Per-hour figures are computed from those quoted rates; add-on features (diarization, summarization, PII redaction) are billed separately.