GPT-6 Astra's Price Sheet: $50 Output and a Trapdoor at 272K
GPT-6 Astra hit OpenRouter on September 4 (https://openrouter.ai/openai/gpt-6-astra), and the listing is the first neutral place to read what OpenAI actually thinks its new model is worth: $10 per million input tokens, $50 per million output, $1 per million cached input, 1.05 million tokens of context, 128K max completion. That is roughly 2.5 times the promotional rate GPT-5.6 Sol launched at. While Anthropic cut cache prices and Google ran introductory pricing, OpenAI went the other way: premium, no discounts.
The detail agent builders actually need is the trapdoor. Per multiple pricing trackers, once your input crosses 272K tokens the entire request reprices — to $20 input, $25 cache write, $75 output. Not the marginal tokens. The whole request. A long-horizon agent that lets its context balloon past that line silently pays roughly double for everything, which makes context discipline a budget line, not a style preference. Fast mode doubles both speed and price on top.
Throughput is the other asymmetry: OpenRouter clocks OpenAI's own Fast endpoint at 63 tokens per second while Azure crawls at 16 to 17 — same model, four times the wall clock, which for agent loops means four times the waiting.
OpenAI's defense showed up the same day in Artificial Analysis data: Astra dominates the output-token-efficiency frontier, doing more per token than anything near its intelligence class. That is the whole bet in one sentence — fewer, smarter, more expensive tokens against everyone else's many cheap ones. The launch story is at https://clauday.com/article/88009f16-a1c6-41ea-8cd9-e572390f7d03. Whether an Astra loop is actually cheaper end-to-end than a Fable 5.1 loop is now the most consequential unmeasured number in the industry.
← Back to all articles
The detail agent builders actually need is the trapdoor. Per multiple pricing trackers, once your input crosses 272K tokens the entire request reprices — to $20 input, $25 cache write, $75 output. Not the marginal tokens. The whole request. A long-horizon agent that lets its context balloon past that line silently pays roughly double for everything, which makes context discipline a budget line, not a style preference. Fast mode doubles both speed and price on top.
Throughput is the other asymmetry: OpenRouter clocks OpenAI's own Fast endpoint at 63 tokens per second while Azure crawls at 16 to 17 — same model, four times the wall clock, which for agent loops means four times the waiting.
OpenAI's defense showed up the same day in Artificial Analysis data: Astra dominates the output-token-efficiency frontier, doing more per token than anything near its intelligence class. That is the whole bet in one sentence — fewer, smarter, more expensive tokens against everyone else's many cheap ones. The launch story is at https://clauday.com/article/88009f16-a1c6-41ea-8cd9-e572390f7d03. Whether an Astra loop is actually cheaper end-to-end than a Fable 5.1 loop is now the most consequential unmeasured number in the industry.
Comments