Morning Edition · Tuesday, August 25, 2026Published at 2:26 AM EDT · New York
The reduction is 20 percent on input and 33 percent on output, and it arrives the same week the model family reaches Amazon's Kiro development environment.

OpenAI lowered the price of its flagship reasoning tier. GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens, down from $5 and $30, a cut of 20 percent on input and 33 percent on output. Winbuzzer reports the rate is guaranteed through at least November 21, 2026, and applies to the pay-as-you-go application programming interface (API), Codex credits, and eligible ChatGPT Work plans. Requests above 272,000 input tokens are billed at twice the input rate and 1.5 times the output rate for the entire request, so the discount shrinks precisely for the long-context requests that agent workloads generate most.
The length of the commitment matters more than the size of the cut. A three-month guaranteed price floor lets a buyer plan unit economics for a full quarter, which is what enterprise procurement actually needs, and it is harder for OpenAI to reverse than a short-term promotional discount. Enterprise DNA counts this as the third price reduction on the GPT-5.6 family within a month.
Distribution expanded at the same time. OpenAI said GPT-5.6 is now available inside Kiro, Amazon's agentic development environment, across three tiers rolling out in Amazon Web Services (AWS) regions in US-East-1 and Europe (Frankfurt). By OpenAI's own measurement, Sol scores 80 on the Coding Agent Index and 88.8 percent on Terminal-Bench 2.1, which the company says beats Claude Fable 5 on both benchmarks while using less than half the output tokens. Terra scores 77.4 and is priced at $2 and $12 per million input and output tokens. Luna scores 74.6 and is priced at $0.20 and $1.20 per million input and output tokens.
Those benchmark figures come from OpenAI and have not been independently verified. The more checkable claim involves token accounting: if Sol genuinely finishes agent tasks using fewer output tokens, its real cost advantage is larger than the listed price suggests, because output tokens make up most of agent billing.
What this means
OpenAI is now competing on committed price and on tokens consumed per completed task at the same time, which pressures rivals along two dimensions at once. Anthropic is the most directly exposed, since coding agents are its strongest commercial position and OpenAI is explicitly benchmarking against Claude Fable 5. Inference margins across the API market compress further as a result, and the vendors best able to absorb that are the ones that control their own serving hardware economics. For engineering teams, the practical shift is that reasoning-tier pricing is now negotiable and time-limited, so architecture decisions that assume a fixed cost per token for a full year are already outdated.
Part of a tracked trend
Frontier Model Price War
Frontier API vendors increasingly compete on strategic price-cutting rather than pure capability, repeatedly launching or repricing models below prevailing rates to grab share and compressing industry inference margins; expect recurring below-rival pricing moves.
Start a discussion in Townsquare.
More from this edition
What to watch
Observations to monitor, not financial advice.
Synthesized from: Polylog editors · OpenAI News
Comments
0No comments yet.