Polylog
The Polylog AI Intelligence Brief

Morning Edition · Friday, July 31, 2026Published at 1:48 AM EDT · New York

OpenAI Cuts GPT-5.6 Luna Prices About 80 Percent, Turning the Model Race Into a Price War

Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, priced below Google's comparable Gemini tier and pressuring rivals' inference margins.

OpenAI Cuts GPT-5.6 Luna Prices About 80 Percent, Turning the Model Race Into a Price War

OpenAI on July 30 lowered the price of its GPT-5.6 line. It cut the lightweight Luna tier by roughly 80 percent to $0.20 per million input tokens and $1.20 per million output tokens, and the mid-tier Terra by about 20 percent to $2 and $12. The top Sol tier keeps its price but gains a new Fast Mode that, according to Russian-language coverage of the release, runs up to 2.5 times faster at double the standard cost, an effective gain of roughly 25 percent per unit of speed.

The company attributes the cuts to efficiency gains realized while building GPT-5.6, including using the model to rewrite and optimize its own serving code. Luna's combined token price of about $1.40 per million is below Google's Gemini 3.5 Flash-Lite at $2.80, while Terra's combined $14 is close to Gemini 3.1 Pro Preview. The framing matters here. OpenAI is presenting this as a shift along the price-performance frontier, not a gain in capability, and the most notable benchmark results of the week came from competitors.

What is verified is the pricing. What is asserted, and worth treating with skepticism, is that these cuts reflect pure efficiency rather than a defensive response to cheaper Chinese open-weight models and to Meta's Muse Spark application programming interface (API), which launched at a claimed 75 percent below rivals. The overall trend is clear either way. The marginal price of a served token at a given capability tier keeps falling.

Veracity: Corroborated
87/100
If true, who benefits

Enterprise API buyers and OpenAI's distribution gain, while whoever holds the lowest serving cost at a given quality tier captures metered demand and defends share against cheaper Chinese open-weight models.

The nuance

The price cut and exact token figures are confirmed by CNBC, VentureBeat and others, but OpenAI's "pure efficiency" attribution omits the competitive timing, with CNBC reporting Chinese models had captured roughly 46 percent of US enterprise token usage on OpenRouter and Meta's discounted API launching the same month.

An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.

What this means

The competitive axis for commodity and mid-tier inference is shifting from capability to cost per token, which compresses gross margins for every metered API vendor at once. Whoever holds the lowest serving cost at a given quality tier gains distribution, and that advantage now comes from inference-stack efficiency and self-optimizing serving code rather than a larger base model. OpenAI, Google, and Anthropic are all exposed through the same channel. Enterprise buyers can now swap providers on price with less capability penalty than a year ago.

What to watch

  • Whether Google and Anthropic match Luna's cut within weeks, which would confirm that mid-tier pricing has become reactive rather than strategic.
  • Reported token volumes after the cut, since an 80 percent price drop only helps revenue if it generates materially more usage.

Observations to monitor, not financial advice.

3 sources

Synthesized from: OpenAI · VentureBeat · Polylog editors

Part of a tracked trend

The Inference-Cost Efficiency Race

Techniques that cut tokens generated and KV-cache memory per query will keep compressing the marginal cost of serving reasoning models, making inference efficiency a recurring competitive axis alongside raw capability.