Gemini ≈ 258 tok/s video, 32 tok/s audio
cost per video
input tokens / video
input tokens / minute
monthly cost

Frame rate is the dial — sweeping it

Image tokens scale directly with frames sampled. For slow footage — lectures, screen recordings, security feeds — a lower frame rate rarely loses anything and cuts the bill in proportion. Audio and prompt are held at your settings.

SamplingInput tokens / videoCost / videoMonthly

The modality nobody budgets for

Text is cheap, images are pricier, and video is in a category of its own — because a model turns a video into a wall of image tokens, one sampled frame at a time, and stacks the audio track on top. The intuition trap is thinking of a video as "one input" like a document; in reality a ten-minute clip at one frame per second is 600 frames, roughly 155,000 image tokens plus 19,000 audio tokens, and a full hour crosses a million input tokens before the model produces a single word of output. That is why a video-understanding feature that looked trivial in a demo can dominate an entire AI bill once it runs across a library. The levers are direct: sample fewer frames per second, drop the audio when you only need the visuals, transcribe speech with a cheap speech-to-text model and reason over the text instead of paying audio-token rates, and clip to the segment that matters. Size the transcript alternative on the speech-to-text cost calculator, the still-image side on the vision image input calculator, and the generation direction on the video generation cost calculator.

Host your project:DigitalOcean — $200 free ↗Hostinger VPS
Vision Image Input CostSpeech-to-Text CostVideo Generation CostGemini API CostVision API Cost