Idle GPUs are the hidden cost
Renting a GPU is cheap per hour and expensive per idle hour. It beats APIs only when utilization is high and steady. Compare against per-token pricing on the self-hosted vs API calculator.
GPU rental: utilization is the real price
Cloud H100s rent for $2โ$10/hour depending on vendor tier โ hyperscalers at the top, bare-metal marketplaces at the bottom. But the quoted rate matters less than your utilization: a $3/hr GPU running training jobs 20% of the time costs $15 per productive hour. Spot/interruptible pricing cuts 50โ70% for checkpointable workloads. The build-vs-rent crossover: a $30k GPU server with 3-year depreciation beats $2.50/hr rentals at roughly 40% sustained utilization โ before counting the ops burden. For inference specifically, per-token serverless GPU offerings usually beat raw rentals below a few million requests monthly.