HomeToolsLearn › GPU Cloud Rental Cost Calculator
per month
$/GPU-hour
spot saves vs on-demand
always-on (730h)

Every GPU at your hours & count

Monthly cost for the hours and GPU count above — on-demand and spot side by side.

GPUVRAMOn-demand $/hSpot $/hMonthly (your tier)
⚠️ Estimate. GPU rates are reference mid-market figures (July 2026) blended across providers like RunPod, Lambda, Vast.ai, CoreWeave and the hyperscalers — actual prices swing by region, availability and commitment. Spot/community capacity can be reclaimed at any time. Always confirm live pricing with your provider. · Report outdated price →

A rented GPU bills every hour it's switched on

The thing that surprises teams moving from APIs to their own GPUs is that the meter never pauses. A per-token API costs nothing when traffic is quiet; a rented H100 costs the same whether it's saturated or sitting at 2% utilization, because you're paying for the hour, not the work. That's why the single biggest lever on your bill isn't the GPU you pick — it's how many hours you keep it running. An H100 you only need eight hours a day costs a third of one left on around the clock. This calculator makes that explicit: set the real hours, not the theoretical maximum, and watch the monthly number move. The quick-set dropdown covers the common patterns — always-on inference, a working-day training box, weekend batch runs.

Tier is the second lever. On-demand capacity is guaranteed for as long as you hold it and priced at a premium; spot, community and interruptible instances are spare capacity sold 20–50% cheaper, with the catch that the provider can take them back. For checkpointed training or fault-tolerant batch jobs, spot is close to free money; for an inference endpoint that can't blink, the on-demand premium buys you stability. The table shows both rates for every card so the trade-off is in front of you. Once you have the rental number, the real question is whether owning GPU-hours actually beats paying per token — work that out with the self-hosted LLM vs API calculator, then price the bandwidth your model serves with the cloud egress calculator and the inference endpoints themselves with the serverless function cost calculator.

How to use it

1. Pick your GPU — its typical on-demand and spot rates load automatically.
2. Choose on-demand or spot, and set how many GPUs you're running.
3. Set the hours per month, or use the quick-set dropdown for a common duty cycle.
4. Read the monthly cost and spot saving, and compare every GPU in the table.

Common mistakes

Assuming always-on. 730 hours is the maximum, not your usage — most workloads need far fewer. Forgetting idle time bills. A GPU left running between jobs still costs full rate; shut it down or use autoscaling. Picking the biggest card. An L4 or 4090 often serves a small model at a fraction of an H100's cost — match the GPU to the model. Ignoring storage and egress. Volumes, snapshots and data transfer are separate line items; add storage above and price egress separately.

Estimate only. GPU rates move constantly and vary by provider and region — verify live pricing before relying on them.

Share: 𝕏 Post Reddit
Host your project:DigitalOcean — $200 free ↗Hostinger VPS
Compare 300+ AI Model PricesLLM Latency & Response Time CalculatorVector DB Storage Cost CalculatorWords to Tokens CalculatorVision API Cost Calculator