Morning Edition · Wednesday, July 22, 2026Published at 1:47 AM EDT · New York
Gemini 3.6 Flash uses 17% fewer output tokens than its predecessor, while 3.5 Flash Cyber, tuned to find and fix vulnerabilities, is restricted to a limited pilot.

Google introduced three new Gemini models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The key figure for engineers is efficiency. On the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash while improving on coding and multimodal tasks. Flash-Lite is positioned for high-throughput agentic work at 350 output tokens per second, priced at $0.30 per million input tokens and $2.50 per million output tokens.
The more consequential release is 3.5 Flash Cyber, a variant fine-tuned to find and fix security vulnerabilities that Google is making available only to governments and trusted partners under a limited-access pilot. Notably, there is no new Pro model in this release, which suggests Google is optimizing the cost-sensitive middle of its lineup rather than pushing the capability ceiling this cycle.
Fewer output tokens at equal quality is a direct cut to inference cost for anyone running Gemini in repeated agentic calls, because reasoning models bill on the tokens they generate. The gated cyber model is the notable governance move. A frontier vendor is treating an offensive-capable security tool as something to distribute by allowlist, the same approach that led to OpenAI's incident report the same day.
What this means
The efficiency gain directly reduces the marginal cost of agentic workloads, where token count, not latency, is the binding cost. Developers building on the Flash tier gain roughly a fifth cheaper generation without switching models, which pressures competing mid-tier APIs from Anthropic and OpenAI on price. The gated Cyber model signals that dual-use security capability is becoming a policy-managed product line rather than an open API, a template other labs are likely to copy.
What to watch
Part of a tracked trend
Frontier Model Efficiency Gains
Capability per unit of training and inference compute keeps improving, letting newer models match prior frontier performance far more cheaply and gradually loosening the link between raw scale and capability.
Start a discussion in Townsquare.
More from this edition
Observations to monitor, not financial advice.
Synthesized from: Google DeepMind · MarkTechPost · TechCrunch
Comments
0No comments yet.