β€”
$ / 1M tokens
β€”tokens / hour
β€”tokens / GPU-month
β€”$ / GPU-month

Idle GPU time is the real inference cost

Cost per token is GPU rent divided by tokens actually served β€” so utilisation, not raw speed, sets your bill. Decide if self-hosting beats a provider on the self-host vs API calculator, and price the underlying hardware on the GPU cloud cost calculator.

Host your project:DigitalOcean β€” $200 free β†—Hostinger VPS
PaymentsEmail & SMSMarket DataMapsSearch & Scraping