Home › Blog

Blog

Numbers-first articles on API pricing, free tiers and cutting your costs.

GPT-5 API pricing: what it actually costs per call (and why it's cheaper than GPT-4o)

2026-07-24 · AI & LLMs
Real 2026 prices with worked per-call math: GPT-5 flagship is $1.25/$10 per 1M — half GPT-4o's input at the same output, with 3× the context. GPT-5 mini and nano compared, and where the money actually goes.
Read →

The tokens are 1% of the bill: what a voice agent that browses and runs code actually costs

2026-07-23 · AI & agents
We priced Cerebras, E2B, Browserbase and LiveKit against one realistic 10-minute agent call. LLM tokens came to 0.89% of the bill; proxy bandwidth came to 74%. Two config habits separate $485 per 1,000 calls from $3,026.
Read →

Cheapest AI API for 20 real workloads — July 2026

2026-07-13 · AI & LLMs
We calculated the cost of 13 AI models across 20 production workloads: customer support, RAG chatbots, content moderation, code review, legal extraction, and more. Gemini Flash-8B wins all 20 on price — 4× cheaper than GPT-4o mini, 145× cheaper than Claude Opus 4.8 for content moderation at 1M calls/month.
Read →

The true first-year cost of Stripe, Twilio and Pinecone

2026-07-05 · Cost strategy
Advertised: ~$3,100. Real year-1 bill: $6,170. We modelled a 500-user SaaS and added every hidden fee — Stripe Radar, Billing add-on, Twilio A2P 10DLC registration, Pinecone storage at scale. The combined stack costs 2× what the pricing pages suggest.
Read →

OpenAI vs Anthropic vs Gemini after prompt caching: real cost difference (2026)

2026-07-15 · AI & LLMs
Caching changes the ranking. Claude charges a write fee but gives 90% off reads. GPT-4o caches automatically with 50% off. Gemini's context cache costs per hour stored. For a 10k-token system prompt at 500 req/day, Claude Sonnet becomes cheapest after caching.
Read →

APIs that quietly raised prices (2024–2026 tracker)

2026-07-09 · Cost strategy
Stripe went from 2.9% to 3.1%. Google Maps doubled some routes. AWS started charging for public IPv4 that used to be free. A running log of provider price increases with old vs new numbers and bill impact.
Read →

What 100,000 customer support chats cost with different LLMs

2026-07-11 · AI & LLMs
Real token math for 100k AI support chats/month. Gemini Flash: $432. GPT-4o: $10,125. The 23× difference, how conversation history grows your bill, and three levers to cut costs without changing models.
Read →

Cheapest LLM for 20 real workloads: actual token math (2026)

2026-07-07 · AI & LLMs
We ran the cost for 20 production workloads — email classification, code review, RAG, legal extraction and more — at 1,000 req/day. The cheapest model is 97% less than frontier for identical classification tasks. Full tables.
Read →

Your AI budget is probably wrong — LLM prices fall ~10× a year

2026-07-11 · Cost strategy
We tracked every major LLM price event since GPT-4's $30 launch: frontier-class input is down 97% in three years, both discontinuous shocks came from DeepSeek, and output prices quietly deflate slower than input. What the ~10×/year curve does to annual contracts, unit economics and build-vs-wait calls.
Read →

Why message 30 in a chatbot costs nearly 7× more than message 1

2026-07-08 · AI & LLMs
Every API call re-sends the full conversation history as input. By turn 30, you're paying to process 11,300 tokens instead of 1,150. A 30-turn conversation totals over $1 — 4× what it would cost stateless. The math behind the accumulation and four ways to manage it.
Read →

The API price is the smallest line: what integrating an API really costs

2026-07-04 · Cost strategy
The per-call rate you compare between vendors is usually under a fifth of the real cost. We added up developer build time, monthly maintenance and usage for a typical integration — and the "cheaper" API on paper is often the pricier one in your ledger. Where volume flips the story, and how to actually compare two APIs.
Read →

Free tier runway — when your $0 API bill becomes a real one

2026-07-02 · Cost strategy
A free tier isn't a plan, it's a runway. At 20% monthly growth, jumping from 40% to 80% utilisation doesn't halve your runway — it cuts it four-fold. We ran the logarithmic math on real 2026 free tiers (Gemini, Resend, Supabase, Cloudflare) and the reset-vs-hard-stop trap that breaks apps in production.
Read →

Grok vs GPT vs Claude vs Gemini — 2026 LLM API pricing, ranked

2026-06-30 · AI & LLMs
Same workload, every major model: $450 on Grok 4.1 Fast versus $17,500 on Claude Opus 4.6 — a 39× spread. xAI took the cheap end of the frontier in 2026, OpenAI and Google split into budget and premium tiers, and Claude stayed top of the ladder. The real per-million numbers.
Read →

Why your LLM app throws 429 errors — the TPM math nobody explains

2026-06-28 · AI & LLMs
Your tier allows 500 requests a minute, so why the 429s at eight? The tokens-per-minute limit binds first — at 4,500 tokens a request, OpenAI's entry tier caps you under 7 RPM, about three concurrent users. The real math on every tier, and the four levers that fix it.
Read →

What a vector database really costs: Pinecone vs Qdrant vs Zilliz vs pgvector

2026-06-25 · AI & LLMs
We size a real RAG workload — 500k vectors, 1536 dimensions, 1M queries/month — and the bill lands in single-to-low-double-digit dollars on every provider. The storage layer is the cheap part of RAG; here's the math and which model to pick.
Read →

What it really costs to send 100,000 SMS: Twilio vs Plivo vs Vonage vs MessageBird

2026-06-23 · Email & SMS
Plivo wins on platform rate, but the US carrier pass-through fee adds ~$300 the same for everyone — and emoji-laden templates quietly bill as 2–3 segments. The all-in math on 100k messages, and when to skip SMS entirely.
Read →

How prompt caching cuts your LLM bill by up to 90%

2026-06-21 · AI & LLMs
Reusing the same system prompt or RAG context? You're paying full price to re-read it. Caching cuts input cost 50–90% — with the 2026 discounts for OpenAI, Anthropic and Gemini, plus the write-premium gotcha that can make your bill go up.
Read →

What a stock market data API really costs: Alpha Vantage vs Finnhub vs Polygon.io

2026-06-18 · Market data
Free tiers compared on calls/min and data freshness, where paid plans start, and the per-asset-class billing quirk that turns Polygon's $29 into $58. Which to pick for dashboards, live quotes or tick data.
Read →

What it really costs to accept payments: Stripe vs PayPal vs Paddle on $10k/mo

2026-06-16 · Payments
A worked breakdown on $10,000/month: why the spread is $3,000/yr, why small orders pay nearly double, and when a Merchant of Record's higher fee is actually the cheaper choice.
Read →

The hidden API costs that quietly double your bill

2026-06-14 · API costs
The headline price is rarely what you pay. SMS carrier fees, email IPs, maps SKUs, payout charges and AI output tokens — the line items that add 30–50% after launch, with real numbers.
Read →

The cheapest AI API in 2026 — real cost at 1 million requests

2026-06-12 · AI & LLMs
Same workload — 1M requests, 1,000 in / 500 out tokens — priced on every major model. The gap between cheapest and priciest is 35×.
Read →

GPT vs Claude vs Gemini: which AI API is cheapest in 2026?

2026-06-10 · AI & LLMs
A numbers-first comparison of OpenAI, Anthropic and Google API pricing — and the 175× spread between cheapest and priciest for the same workload.
Read →

The best free API tiers in 2026

2026-06-02 · Roundup
Genuinely useful free API tiers across AI, email, maps, crypto data and search — and how far each really gets you.
Read →

Stripe vs PayPal vs Paddle: payment fees compared

2026-05-23 · Payments
Payment API fees compared — and when a merchant-of-record like Paddle is worth the higher cut.
Read →

Twilio vs Vonage vs MessageBird: SMS API cost compared

2026-05-15 · SMS
Why the carrier fee, not the platform, often decides your SMS bill.
Read →

The cheapest way to send transactional email in 2026

2026-05-07 · Email
SendGrid vs Resend vs Amazon SES vs Postmark — pricing, free tiers and which is cheapest at your volume.
Read →

OpenAI vs Anthropic Claude: which API should you build on?

2026-04-29 · AI & LLMs
Pricing, models, strengths — and which to pick for your use case.
Read →