# OpenAI Cuts GPT-5.6 Luna Prices About 80 Percent, Turning the Model Race Into a Price War

Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, priced below Google's comparable Gemini tier and pressuring rivals' inference margins.

- Published: 2026-07-31T05:48:01.029Z
- Canonical: https://polylog.news/ai/2026-07-31/openai-cuts-gpt-5-6-luna-prices-about-80-percent-turning-the
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [OpenAI](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6), [VentureBeat](https://venturebeat.com/technology/ai-price-wars-openai-cuts-gpt-5-6-luna-prices-by-80-as-model-competition-shifts-toward-cost), [Polylog editors](https://polylog.news)

OpenAI on July 30 lowered the price of its GPT-5.6 line. It cut the lightweight [Luna tier by roughly 80 percent](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6) to $0.20 per million input tokens and $1.20 per million output tokens, and the mid-tier Terra by about 20 percent to $2 and $12. The top Sol tier keeps its price but gains a new Fast Mode that, [according to Russian-language coverage of the release](https://t.me/ai_machinelearning_big_data/10622), runs up to 2.5 times faster at double the standard cost, an effective gain of roughly 25 percent per unit of speed.

The company attributes the cuts to efficiency gains realized while building GPT-5.6, including using the model to rewrite and optimize its own serving code. Luna's combined token price of about $1.40 per million is [below Google's Gemini 3.5 Flash-Lite at $2.80](https://venturebeat.com/technology/ai-price-wars-openai-cuts-gpt-5-6-luna-prices-by-80-as-model-competition-shifts-toward-cost), while Terra's combined $14 is close to Gemini 3.1 Pro Preview. The framing matters here. OpenAI is presenting this as a shift along the price-performance frontier, not a gain in capability, and the most notable benchmark results of the week came from competitors.

What is verified is the pricing. What is asserted, and worth treating with skepticism, is that these cuts reflect pure efficiency rather than a defensive response to cheaper Chinese open-weight models and to Meta's Muse Spark application programming interface (API), which launched at a claimed 75 percent below rivals. The overall trend is clear either way. The marginal price of a served token at a given capability tier keeps falling.

## What this means

The competitive axis for commodity and mid-tier inference is shifting from capability to cost per token, which compresses gross margins for every metered API vendor at once. Whoever holds the lowest serving cost at a given quality tier gains distribution, and that advantage now comes from inference-stack efficiency and self-optimizing serving code rather than a larger base model. OpenAI, Google, and Anthropic are all exposed through the same channel. Enterprise buyers can now swap providers on price with less capability penalty than a year ago.

## What to watch

- Whether Google and Anthropic match Luna's cut within weeks, which would confirm that mid-tier pricing has become reactive rather than strategic.
- Reported token volumes after the cut, since an 80 percent price drop only helps revenue if it generates materially more usage.
