Morning Edition · Thursday, September 3, 2026Published at 2:26 AM EDT · New York
The introductory rate runs only through December 31, after which Google doubles it, making the next four months a deliberate push to win market share in the cheap tier.

Google released Gemini 3.8 Flash and a security-tuned sibling, 3.8 Flash Cyber, three weeks after its previous Flash refresh, according to 9to5Google. The model keeps a roughly one-million-token context window and ships at $0.75 per million input tokens and $3.75 per million output tokens, introductory pricing that expires on December 31, 2026 and then rises to $1.50 and $7.50.
The headline number is the one to treat carefully. On the verified subset of Humanity's Last Exam, a multi-step reasoning test spanning science, humanities and professional questions, Gemini 3.8 Flash scored 54.9 percent against 54.4 percent for Claude Opus 5, as reported by the Telegram channel AI Post. A half-point gap on a single evaluation is not a capability difference. It is a tie that happens to fall on the favorable side for the vendor publishing it, and no independent reproduction has appeared yet.
The agentic coding claim is larger and more version-dependent. Google reports 90.8 percent on Terminal-Bench 2.1, a benchmark for agentic command-line tasks, up from 81.6 percent for Gemini 3.7 Flash. On the newer and harder Terminal-Bench 4.0, a Hacker News commenter put the same model at 19.1 percent against Opus 5 at 51.8 percent. Both figures can be true at once, which is the point: which benchmark version a lab chooses to publish now matters as much as the model itself.
The Russian-language channel AI ML Big Data described the release as optimized for large mixed text-and-image workloads with a strong emphasis on software development, which matches how Google is positioning the Flash tier. The strategic content of the launch is not the reasoning score. It is a frontier lab putting near-frontier reasoning into its cheapest served tier, at a price that undercuts what premium models charge for comparable work, and putting a date on when that discount ends.
Part of a tracked trend
Frontier Model Price War
Frontier API vendors increasingly compete on strategic price-cutting rather than pure capability, repeatedly launching or repricing models below prevailing rates to grab share and compressing industry inference margins; expect recurring below-rival pricing moves.
Start a discussion in Townsquare.
More from this edition
Google, which converts a half-point benchmark tie into an "Opus-class" marketing claim while buying fourth-quarter workload migration at a price it has already scheduled to double, pressuring the inference margins of OpenAI and Anthropic.
The $0.75 price and the 54.9 percent Humanity's Last Exam verified-subset score check out as reported, but both come from Google's own runs, and the weak Terminal-Bench 4.0 result the article attributes to a Hacker News commenter appears in Google's own published comparison chart, which changes who is disclosing the unflattering number.
An open-source-intelligence read of how likely this story is true with its real nuance, not a judgment of any outlet. It assesses the claim, weighing independent and adversarial reporting. How we label confidence.
What this means
Google is competing on cost per solved task rather than on a capability lead it does not clearly have. Anyone selling reasoning tokens at premium-tier prices, including OpenAI's and Anthropic's flagship endpoints, now has to justify the gap on evaluations where the margin is larger than half a point. The December 31 expiry reveals the strategy: Google is buying workload migration during a window when switching costs are low, betting that agent pipelines built on Flash in the fourth quarter will stay there once the price doubles in January. Inference gross margins across the API market compress either way, because rivals rarely respond to a competitor pricing at half the going rate by raising their own prices.
What to watch
Observations to monitor, not financial advice.
Synthesized from: Google DeepMind · Polylog editors · Hacker News
Comments
0No comments yet.