使用状況とクラウドプライシング
デバイスの機能とフォールバック
コストの内訳
| 品目 | 推論/月 | 月あたりのコスト |
|---|---|---|
| 純粋なクラウドのベースライン (100% クラウド、オンデバイスなし) | — | — |
| — クラウド専用ユーザー (オンデバイス対応ではない) | — | — |
| — オンデバイス対応ユーザーからのクラウド フォールバック | — | — |
| = ハイブリッド クラウドの請求書 | — | — |
| + 償却後のオンデバイスセットアップコスト | — | — |
| = ハイブリッドの月額総コスト | — | — |
これが他のツールにどのように接続されるか
このサイトにはすでに 2 つの計算ツールがあり、関連するものの構造的に異なる決定の価格を決定しています。の セルフホストと API のコスト計算ツール そして セルフホスト型 LLM と API の計算ツール 両方のモデル サーバー側 セルフホスティング — トークンごとまたは呼び出しごとにプロバイダーに支払うのと比較して、独自の GPU またはクラウド インスタンスをレンタルし、そこで推論を自分で実行します。どちらのコスト形態でも、実際のホスティング料金がかかるサーバーはどこかにあり、ベンダーのものではなくあなたが管理するものにすぎません。この計算機は代わりにモデルを作成します クライアント側 オンデバイス推論はエンド ユーザーの電話ハードウェア上で直接実行されます。サーバーをレンタルする必要はなく、モデルの出荷後は推論あたりの限界コストが事実上ゼロになります。また、まったく別のボトルネックがあります。GPU 時間ではなく、十分な機能を備えたデバイスを所有しているユーザーの割合と、それらのデバイスでさえクラウドにフォールバックする必要がある頻度です。 GPU ボックスをレンタルするか、クラウド API に料金を支払うかを決定する場合は、セルフホスト型の計算ツールを使用してください。量子化モデルをモバイル アプリ自体の内部に出荷するかどうかを決定している場合、これがそのコスト形状に一致するものになります。
数字の読み方
At the defaults — 100,000 monthly active users, 50 inferences per user per day, $0.001 per cloud inference, a $40,000 one-time on-device build, 60% of users on capable devices, a 5% cloud-fallback rate on those capable users, amortized over 12 months — the pure-cloud baseline runs $150,000/month. Going hybrid brings the residual cloud bill down to $64,500/month (60,000,000 inferences from cloud-only users plus 4,500,000 fallback inferences from capable users), adds roughly $3,333/month in amortized setup cost, and lands on a hybrid total near $67,833/month — a savings of about $82,167/month, or roughly 純粋なクラウドベースラインから55%オフ. The $40,000 setup cost itself breaks even in under half a month at this volume, because the monthly cloud spend it deflects (about $85,500/month) dwarfs the one-time build cost almost immediately. That math changes fast at lower scale: run the same inputs at 5,000 users instead of 100,000 and the break-even stretches to roughly 20x longer — about 9.4 months instead of under half a month — because there simply are not enough deflected inferences per month to recoup the build cost quickly. Plug in your app's real MAU, not an aspirational one, before committing engineering budget to an on-device build.