Morning Edition · Friday, July 31, 2026Published at 2:00 AM EDT · New York
OpenAI Cuts GPT-5.6 Luna Pricing 80 Percent and Adds a Faster Sol Mode
Luna drops to 0.20 dollars per million input tokens and 1.20 dollars output, while a new Sol Fast mode runs up to 2.5 times faster at twice the standard price.

OpenAI repriced its GPT-5.6 tiers on July 30, and the change is about cost, not capability. Luna, the cheapest tier, falls 80 percent to 0.20 dollars per million input tokens and 1.20 dollars per million output tokens. Terra, the balanced tier, drops about 20 percent to 2 and 12 dollars. The flagship Sol tier keeps standard pricing at 5 and 30 dollars but gains a Fast mode that delivers up to 2.5 times the throughput at twice the price, 10 and 60 dollars per million tokens, with no change in intelligence.
The competitive implication is specific. At a combined 1.40 dollars per million tokens, Luna now undercuts Google's Gemini 3.5 Flash-Lite on the high-throughput, low-latency workloads where per-request cost compounds: classification, routing, summarization, and lightweight real-time assistants. Independent tracking placed Luna's cost-per-task 80 percent lower after the cut. By the same analysis, Terra still ranks behind Luna and Sol on the cost-performance frontier despite its reduction.
OpenAI is describing efficiency gains that it passes through as lower prices rather than higher benchmark scores. Fast mode is a serving optimization sold as a premium, while the cheap tiers absorb the falling marginal cost of inference. For engineers, the actionable change is at the routing layer. A task that was economically borderline on Luna last week is roughly five times cheaper this week, which shifts where it makes sense to spend a frontier call versus a small-model call.
What this means
The axis of competition among frontier vendors is moving from raw capability toward cost per task, and the mechanism is inference efficiency converted directly into application programming interface (API) pricing. OpenAI gains distribution on high-volume workloads and pressures Google's Flash-Lite and Gemini Flash pricing, while the losers are any provider whose small-model tier is priced above the new floor. For buyers, the cost of switching cheap tasks to the lowest bidder just dropped.
What to watch
- Whether Google and Anthropic respond with matching cuts to their cheap tiers, which would confirm a sustained price war rather than a one-off.
- Independent cost-per-task benchmarks (not vendor token prices) that reveal whether Luna's real economics beat rivals once output-token verbosity is counted.
Observations to monitor, not financial advice.
Synthesized from: OpenAI · VentureBeat · Polylog editors
Part of a tracked trend
The Inference-Cost Efficiency Race
Techniques that cut tokens generated and KV-cache memory per query will keep compressing the marginal cost of serving reasoning models, making inference efficiency a recurring competitive axis alongside raw capability.
More from this edition
- Bond Investors Start Repricing the Debt Financing the AI Buildout
- US Regulator Bans New Imports of Foreign-Made Humanoid and Quadruped Robots
- DeepMind's Gemini Robotics ER 2 Adds Video Progress Tracking and Multi-Robot Coordination
- DeepMind Reassigns Its AlphaFold Team, Redirecting Talent Toward Gemini
- Zuckerberg Urges Washington to Accelerate AI Rather Than Restrict It
- GPTZero Finds Fabricated Citations in PwC Middle East Research Reports
- Paper Documents Emergent Deception in Mixed-Motive LLM Multi-Agent Systems
- New Jailbreak Method Uses Dual-Layer Encoding to Slip Past LLM Moderation
- Italian Startup Unveils a Humanoid With Full-Body Sensor Skin
- ByteDance Readies Seedance 2.5, Its Next AI Video Generation Model
- Study Probes Why RL-Tuned Models Out-Reason Their Supervised Counterparts