# Google Ships a Token-Cheaper Gemini Flash Tier and a Cyber-Specialized Model Gated to Governments

Gemini 3.6 Flash uses 17% fewer output tokens than its predecessor, while 3.5 Flash Cyber, tuned to find and fix vulnerabilities, is restricted to a limited pilot.

- Published: 2026-07-22T05:47:02.310Z
- Canonical: https://polylog.news/ai/2026-07-22/google-ships-a-token-cheaper-gemini-flash-tier-and-a-cyber-s
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [Google DeepMind](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/), [MarkTechPost](https://www.marktechpost.com/2026/07/21/google-releases-gemini-3-6-flash-3-5-flash-lite-and-3-5-flash-cyber-a-cheaper-more-token-efficient-flash-tier-built-for-agentic-workloads/), [TechCrunch](https://techcrunch.com/2026/07/21/google-releases-three-new-gemini-models-but-no-3-5-pro/)

Google [introduced three new Gemini models](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/): Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The key figure for engineers is efficiency. On the Artificial Analysis Index, 3.6 Flash [uses 17% fewer output tokens](https://www.marktechpost.com/2026/07/21/google-releases-gemini-3-6-flash-3-5-flash-lite-and-3-5-flash-cyber-a-cheaper-more-token-efficient-flash-tier-built-for-agentic-workloads/) than 3.5 Flash while improving on coding and multimodal tasks. Flash-Lite is positioned for high-throughput agentic work at 350 output tokens per second, priced at $0.30 per million input tokens and $2.50 per million output tokens.

The more consequential release is 3.5 Flash Cyber, a variant fine-tuned to find and fix security vulnerabilities that Google is making available [only to governments and trusted partners](https://techcrunch.com/2026/07/21/google-releases-three-new-gemini-models-but-no-3-5-pro/) under a limited-access pilot. Notably, there is no new Pro model in this release, which suggests Google is optimizing the cost-sensitive middle of its lineup rather than pushing the capability ceiling this cycle.

Fewer output tokens at equal quality is a direct cut to inference cost for anyone running Gemini in repeated agentic calls, because reasoning models bill on the tokens they generate. The gated cyber model is the notable governance move. A frontier vendor is treating an offensive-capable security tool as something to distribute by allowlist, the same approach that led to OpenAI's incident report the same day.

## What this means

The efficiency gain directly reduces the marginal cost of agentic workloads, where token count, not latency, is the binding cost. Developers building on the Flash tier gain roughly a fifth cheaper generation without switching models, which pressures competing mid-tier APIs from Anthropic and OpenAI on price. The gated Cyber model signals that dual-use security capability is becoming a policy-managed product line rather than an open API, a template other labs are likely to copy.

## What to watch

- Independent reproduction of the 17% token-reduction claim on Artificial Analysis, since the figure is currently a vendor citation of a third-party index rather than a checked result.
- How Google scopes access to Flash Cyber, which will show whether allowlisting offensive-capable models becomes an industry norm or a Google-specific choice.
