# Google Prices Gemini 3.8 Flash at 75 Cents per Million Input Tokens and Claims Opus-Class Reasoning

The introductory rate runs only through December 31, after which Google doubles it, making the next four months a deliberate push to win market share in the cheap tier.

- Published: 2026-09-03T06:26:17.363Z
- Canonical: https://polylog.news/ai/2026-09-03/google-prices-gemini-3-8-flash-at-75-cents-per-million-input
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Google DeepMind](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/), [Polylog editors](https://polylog.news), [Hacker News](https://news.ycombinator.com/item?id=gemini-3-8-flash)

Google released [Gemini 3.8 Flash and a security-tuned sibling, 3.8 Flash Cyber](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/), three weeks after its previous Flash refresh, according to [9to5Google](https://9to5google.com/2026/09/02/gemini-3-8-flash-launch/). The model keeps a roughly one-million-token context window and ships at $0.75 per million input tokens and $3.75 per million output tokens, [introductory pricing that expires on December 31, 2026](https://tech-insider.org/gemini-3-8-flash-launch-pricing-2026/) and then rises to $1.50 and $7.50.

The headline number is the one to treat carefully. On the verified subset of Humanity's Last Exam, a multi-step reasoning test spanning science, humanities and professional questions, Gemini 3.8 Flash scored 54.9 percent against 54.4 percent for Claude Opus 5, [as reported by the Telegram channel AI Post](https://t.me/aipost/8030). A half-point gap on a single evaluation is not a capability difference. It is a tie that happens to fall on the favorable side for the vendor publishing it, and no independent reproduction has appeared yet.

The agentic coding claim is larger and more version-dependent. Google reports 90.8 percent on Terminal-Bench 2.1, a benchmark for agentic command-line tasks, up from 81.6 percent for Gemini 3.7 Flash. On the newer and harder Terminal-Bench 4.0, a Hacker News commenter put the same model at 19.1 percent against Opus 5 at 51.8 percent. Both figures can be true at once, which is the point: which benchmark version a lab chooses to publish now matters as much as the model itself.

The Russian-language channel AI ML Big Data described the release as [optimized for large mixed text-and-image workloads with a strong emphasis on software development](https://t.me/ai_machinelearning_big_data/10820), which matches how Google is positioning the Flash tier. The strategic content of the launch is not the reasoning score. It is a frontier lab putting near-frontier reasoning into its cheapest served tier, at a price that undercuts what premium models charge for comparable work, and putting a date on when that discount ends.

## What this means

Google is competing on cost per solved task rather than on a capability lead it does not clearly have. Anyone selling reasoning tokens at premium-tier prices, including OpenAI's and Anthropic's flagship endpoints, now has to justify the gap on evaluations where the margin is larger than half a point. The December 31 expiry reveals the strategy: Google is buying workload migration during a window when switching costs are low, betting that agent pipelines built on Flash in the fourth quarter will stay there once the price doubles in January. Inference gross margins across the API market compress either way, because rivals rarely respond to a competitor pricing at half the going rate by raising their own prices.

## What to watch

- Whether independent evaluators such as Artificial Analysis reproduce the Humanity's Last Exam and Terminal-Bench figures, since a vendor-only number that no third party can hit would mark the score as a configuration artifact rather than a capability.
- Whether OpenAI or Anthropic reprice a mid-tier model within weeks, which would confirm that per-token pricing, not benchmark leadership, is now the primary basis of competition.
- How many developers actually migrate production agent workloads before the January price change, because that migration is what Google is really purchasing with the discount.
