Cache audio โ never synthesise twice
TTS bills per character, so re-generating the same phrase is pure waste. Cache and reuse clips, and match the voice tier to the need. For the reverse direction see the speech-to-text calculator.
What generating spoken audio from text costs per month.
TTS bills per character, so re-generating the same phrase is pure waste. Cache and reuse clips, and match the voice tier to the need. For the reverse direction see the speech-to-text calculator.
Free text-to-speech API cost calculator โ characters synthesised per month and per-character price to your TTS bill.
Almost always per character of input text, quoted per 1 million characters. Standard voices are cheapest; high-quality neural, expressive and cloned voices cost several times more per character.
Volume and voice tier. Narrating long content (audiobooks, courses) burns millions of characters, and premium voices multiply the rate. Cache generated audio so you never pay to synthesise the same text twice.
Cost at different volumes:
| Volume | Monthly | Yearly |
|---|