โ€”
$ / 1M tokens
โ€”tokens / hour
โ€”tokens / GPU-month
โ€”$ / GPU-month

Idle GPU time is the real inference cost

Cost per token is GPU rent divided by tokens actually served โ€” so utilisation, not raw speed, sets your bill. Decide if self-hosting beats a provider on the self-host vs API calculator, and price the underlying hardware on the GPU cloud cost calculator.

Host your project:DigitalOcean โ€” $200 free โ†—Hostinger VPS
PaymentsEmail & SMSMarket DataMapsSearch & Scraping