20 September 2026 · AI & LLMs · 6 min read
GPT-6 Astra launched September 3 at $10 per million input tokens and $50 per million output — 2.5x what the previous flagship, GPT-5.6 Sol, charges. That headline number is the one everyone quoted. What most coverage skipped is that $10/$50 is only the middle of three prices for the exact same model, same weights, same output quality. Astra Fast bills $20/$100. Astra Batch bills $5/$25. Pick the wrong one for your workload and you're paying 4x more than you need to for output a human reviewer couldn't tell apart. I pulled OpenAI's live pricing docs, checked them against our own rate-tracking, and ran three real workloads through all three modes to find out where that 4x actually shows up on a bill.
This is new with Astra. GPT-5 and GPT-5.6 priced by model tier — you picked Sol, Terra, or Luna and that set your rate. Astra keeps the tier choice but adds a second axis: mode. Same reasoning depth, same weights, three latency/price points.
| Режим | Input /1M | Output /1M | What it trades |
|---|---|---|---|
| Astra Fast | $20.00 | $100.00 | Lower latency, 2x standard |
| Astra Standard | $10.00 | $50.00 | Default synchronous rate |
| Astra Batch/Flex | $5.00 | $25.00 | Async, 0.5x standard |
| Cached input | $1.00 | — | 0.1x standard input |
| Cache write | $12.50 | — | 1.25x standard input |
| GPT-5.6 Sol (previous flagship) | $5.00 | $30.00 | For comparison |
Fast costs exactly double Standard on both input and output. Batch costs exactly half. That makes Fast four times the price of Batch for output nobody can distinguish — the only difference is whether you got the token stream in under a second or picked it up from a completed job an hour later. There's also a fourth, quieter price: requests carrying more than 272,000 input tokens bill at $20/$75 even on the standard tier, so a long-context workload can slide into Fast-mode pricing without anyone touching the mode selector.
Take a support copilot handling 3,000 live conversations a day, each pulling roughly 6,000 tokens of ticket history and knowledge-base context in, and writing a 400-token reply. That's 90,000 requests a month: 540M input tokens, 36M output tokens.
| Режим | Input (540M) | Output (36M) | Итого за месяц |
|---|---|---|---|
| Astra Fast | $10,800.00 | $3,600.00 | $14,400.00 |
| Astra Standard | $5,400.00 | $1,800.00 | $7,200.00 |
| GPT-5.6 Sol (standard) | $2,700.00 | $1,080.00 | $3,780.00 |
Fast costs $7,200 a month more than Standard for the identical answer, just delivered faster. For a live chat widget, that premium is at least arguable — users notice a slow reply. But notice the Astra-vs-Sol gap too: $7,200 vs $3,780 isn't the full 2.5x you'd expect from the headline rate, it's 1.9x, because output is only 1.67x more expensive on Astra while input is a clean 2x. Your actual premium over Sol depends on how output-heavy your workload is — a terse-reply bot pays closer to 2x, a workload with long completions pays closer to 1.67x.
A coding-agent product that resends an 8,000-token repo context on every follow-up turn, averaging 15 turns per session, at 20,000 sessions a month. Standard mode, no cache: 15 turns × 8,000 tokens × $10/M = $1.20 per session.
| Per session | At 20,000 sessions/mo | |
|---|---|---|
| No caching (120,000 ctx tokens × $10/M) | $1.2000 | $24,000.00 |
| With caching (1 write + 14 hits) | $0.2120 | $4,240.00 |
That's $19,760 a month back — an 82% cut — from turning on prompt caching for context that was already being resent verbatim, no mode change required. It's a bigger saving than switching from Fast to Standard mode entirely. If your agent product is on Astra and isn't caching the repo or conversation context it resends every turn, that's the first thing to fix, before you touch which mode you're calling.
A nightly job tagging and categorizing support tickets: 50M input tokens, 5M output tokens, no human waiting on the result. This is exactly what Batch mode is built for, so I priced it against Fast to see the ceiling-to-floor spread on the same job.
| Режим | Input (50M) | Output (5M) | Итого за месяц |
|---|---|---|---|
| Astra Fast | $1,000.00 | $500.00 | $1,500.00 |
| Astra Standard | $500.00 | $250.00 | $750.00 |
| Astra Batch/Flex | $250.00 | $125.00 | $375.00 |
$1,125 a month, exactly a 4x difference, for a job where nobody is watching a spinner. That's the cleanest case in this whole rate card: there's no product reason a nightly batch tagging run should ever touch Fast or even Standard pricing. If it can run async, route it to Batch and bank the other three-quarters of the bill.
Default new integrations to Standard mode and leave Fast for the specific endpoints where a user is staring at a loading indicator — not for background jobs that happen to call the same model. Route anything asynchronous to Batch without a second thought; the 4x gap over Fast is too large to leave on the table for a job with no latency requirement. And before reaching for a mode change at all, check whether you're resending the same context turn after turn — caching cut this workload's bill more than any mode switch would have. Last, watch your input size on Astra specifically: cross 272K tokens and you're paying Fast-mode rates on the standard tier whether you meant to or not.
Новичок в расценках LLM на основе использования? Начните с бесплатные руководства по стоимости API.
Разверните его самостоятельно: DigitalOcean — бесплатный кредит в размере 200 долларов США ↗ · Хостингер VPS ↗
Pricing verified against OpenAI's official API pricing docs (openai.com/api/pricing) and our own live rate tracking on apicostcalc.com/openai.html, checked 2026-09-20. GPT-6 Astra launched September 3, 2026. Reference estimates using OpenAI's published self-serve rates — Azure OpenAI, Bedrock pricing, enterprise volume discounts and future rate changes vary; confirm current pricing at the source before budgeting. Workload figures (request volumes, token counts) are illustrative scenarios computed against real published per-token rates, not vendor-supplied numbers.