20 settembre 2026 · AI & LLMs · 6 min letto
GPT-6 Astra launched September 3 at $10 per million input tokens and $50 per million output — 2.5x what the previous flagship, GPT-5.6 Sol, charges. That headline number is the one everyone quoted. What most coverage skipped is that $10/$50 is only the middle of three prices for the exact same model, same weights, same output quality. Astra Fast bills $20/$100. Astra Batch bills $5/$25. Pick the wrong one for your workload and you're paying 4x more than you need to for output a human reviewer couldn't tell apart. I pulled OpenAI's live pricing docs, checked them against our own rate-tracking, and ran three real workloads through all three modes to find out where that 4x actually shows up on a bill.
Questa è una novità con Astra. GPT-5 e GPT-5.6 a seconda del livello del modello: hai scelto Sol, Terra o Luna e questo ha determinato la tua tariffa. Astra mantiene la scelta del livello ma aggiunge un secondo asse: la modalità. Stessa profondità di ragionamento, stessi pesi, tre punti di latenza/prezzo.
| Modalità | Input /1M | Uscita /1M | Cosa scambia |
|---|---|---|---|
| Astra Fast | $20.00 | $100.00 | Latenza inferiore, standard 2x |
| Astra Standard | $10.00 | $50.00 | Velocità sincrona predefinita |
| Astra Batch/Flex | $5.00 | $25.00 | Asincrono, standard 0,5x |
| Ingresso memorizzato nella cache | $1.00 | — | immissione standard |
| Scrittura cache | $12.50 | — | 1,25x ingresso standard |
| GPT-5.6 Sol (precedente ammiraglia) | $5.00 | $30.00 | Per il confronto |
Fast costs exactly double Standard on both input and output. Batch costs exactly half. That makes Fast four times the price of Batch for output nobody can distinguish — the only difference is whether you got the token stream in under a second or picked it up from a completed job an hour later. There's also a fourth, quieter price: requests carrying more than 272,000 input tokens bill at $20/$75 even on the standard tier, so a long-context workload can slide into Fast-mode pricing without anyone touching the mode selector.
Prendi un copilota di supporto che gestisce 3.000 conversazioni dal vivo al giorno, ognuna delle quali estrae circa 6.000 token di cronologia dei ticket e contesto della base di conoscenze e scrive una risposta da 400 token. Sono 90.000 richieste al mese: 540 MILIONI di token di input, 36 MILIONI di token di output.
| Modalità | Ingresso (540 M) | Uscita (36 M) | Totale Mensile |
|---|---|---|---|
| Astra Fast | $10,800.00 | $3,600.00 | $14,400.00 |
| Astra Standard | $5,400.00 | $1,800.00 | $7,200.00 |
| GPT-5.6 Sol (standard) | $2,700.00 | $1,080.00 | $3,780.00 |
Fast costs $7,200 a month more than Standard for the identical answer, just delivered faster. For a live chat widget, that premium is at least arguable — users notice a slow reply. But notice the Astra-vs-Sol gap too: $7,200 vs $3,780 isn't the full 2.5x you'd expect from the headline rate, it's 1.9x, because output is only 1.67x more expensive on Astra while input is a clean 2x. Your actual premium over Sol depends on how output-heavy your workload is — a terse-reply bot pays closer to 2x, a workload with long completions pays closer to 1.67x.
Un prodotto agente di codifica che invia nuovamente un contesto repo di 8.000 token a ogni turno di follow-up, con una media di 15 turni per sessione, a 20.000 sessioni al mese. Modalità standard, senza cache: 15 turni × 8.000 gettoni × $10/M = $1,20 per sessione.
| Per sessione | A 20.000 sessioni/mese | |
|---|---|---|
| Nessuna cache (120.000 ctx token × $10/M) | $1.2000 | $24,000.00 |
| Con caching (1 scrittura + 14 hit) | $0.2120 | $4,240.00 |
Sono $19.760 al mese fa — un taglio dell'82% — dall'attivazione della cache dei prompt per il contesto che era già stato risentito alla lettera, non è necessario alcun cambio di modalità. È un risparmio maggiore rispetto al passaggio dalla modalità Fast a quella Standard. Se il tuo prodotto agente è su Astra e non memorizza nella cache il repository o il contesto della conversazione che invia ogni volta, questa è la prima cosa da risolvere, prima di toccare la modalità che stai chiamando.
Un lavoro notturno che tagga e categorizza i ticket di supporto: 50 MILIONI di token di input, 5 milioni di token di output, nessun essere umano in attesa del risultato. Questo è esattamente ciò per cui è stata creata la modalità Batch, quindi l'ho valutata rispetto a Fast per vedere la distribuzione da soffitto a pavimento sullo stesso lavoro.
| Modalità | Ingresso (50M) | Uscita (5M) | Totale Mensile |
|---|---|---|---|
| Astra Fast | $1,000.00 | $500.00 | $1,500.00 |
| Astra Standard | $500.00 | $250.00 | $750.00 |
| Astra Batch/Flex | $250.00 | $125.00 | $375.00 |
$1.125 al mese, esattamente una differenza di 4 volte, per un lavoro in cui nessuno sta guardando uno spinner. Questo è il caso più pulito in questa scheda tariffaria completa: non c'è motivo per cui una sessione di etichettatura batch serale debba mai toccare i prezzi Fast o Standard. Se può essere eseguito in modo asincrono, instradarlo a Batch e incassare gli altri tre quarti della bolletta.
Default new integrations to Standard mode and leave Fast for the specific endpoints where a user is staring at a loading indicator — not for background jobs that happen to call the same model. Route anything asynchronous to Batch without a second thought; the 4x gap over Fast is too large to leave on the table for a job with no latency requirement. And before reaching for a mode change at all, check whether you're resending the same context turn after turn — caching cut this workload's bill more than any mode switch would have. Last, watch your input size on Astra specifically: cross 272K tokens and you're paying Fast-mode rates on the standard tier whether you meant to or not.
Non conosci i prezzi LLM basati sull'utilizzo? Inizia con il guide gratuite sui costi API.
Distribuiscilo tu stesso: DigitalOcean: credito gratuito di $ 200 ↗ · Hostinger VPS ↗
Pricing verified against OpenAI's official API pricing docs (openai.com/api/pricing) and our own live rate tracking on apicostcalc.com/openai.html, checked 2026-09-20. GPT-6 Astra launched September 3, 2026. Reference estimates using OpenAI's published self-serve rates — Azure OpenAI, Bedrock pricing, enterprise volume discounts and future rate changes vary; confirm current pricing at the source before budgeting. Workload figures (request volumes, token counts) are illustrative scenarios computed against real published per-token rates, not vendor-supplied numbers.