Polylog
The Polylog AI Intelligence Brief

Morning Edition · Wednesday, July 22, 2026Published at 1:47 AM EDT · New York

Google Ships a Token-Cheaper Gemini Flash Tier and a Cyber-Specialized Model Gated to Governments

Gemini 3.6 Flash uses 17% fewer output tokens than its predecessor, while 3.5 Flash Cyber, tuned to find and fix vulnerabilities, is restricted to a limited pilot.

Google Ships a Token-Cheaper Gemini Flash Tier and a Cyber-Specialized Model Gated to Governments

Google introduced three new Gemini models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The key figure for engineers is efficiency. On the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash while improving on coding and multimodal tasks. Flash-Lite is positioned for high-throughput agentic work at 350 output tokens per second, priced at $0.30 per million input tokens and $2.50 per million output tokens.

The more consequential release is 3.5 Flash Cyber, a variant fine-tuned to find and fix security vulnerabilities that Google is making available only to governments and trusted partners under a limited-access pilot. Notably, there is no new Pro model in this release, which suggests Google is optimizing the cost-sensitive middle of its lineup rather than pushing the capability ceiling this cycle.

Fewer output tokens at equal quality is a direct cut to inference cost for anyone running Gemini in repeated agentic calls, because reasoning models bill on the tokens they generate. The gated cyber model is the notable governance move. A frontier vendor is treating an offensive-capable security tool as something to distribute by allowlist, the same approach that led to OpenAI's incident report the same day.

What this means

The efficiency gain directly reduces the marginal cost of agentic workloads, where token count, not latency, is the binding cost. Developers building on the Flash tier gain roughly a fifth cheaper generation without switching models, which pressures competing mid-tier APIs from Anthropic and OpenAI on price. The gated Cyber model signals that dual-use security capability is becoming a policy-managed product line rather than an open API, a template other labs are likely to copy.

What to watch

  • Independent reproduction of the 17% token-reduction claim on Artificial Analysis, since the figure is currently a vendor citation of a third-party index rather than a checked result.
  • How Google scopes access to Flash Cyber, which will show whether allowlisting offensive-capable models becomes an industry norm or a Google-specific choice.

Observations to monitor, not financial advice.

3 sources

Synthesized from: Google DeepMind · MarkTechPost · TechCrunch

Part of a tracked trend

Frontier Model Efficiency Gains

Capability per unit of training and inference compute keeps improving, letting newer models match prior frontier performance far more cheaply and gradually loosening the link between raw scale and capability.