Google launched three Flash series models simultaneously in July 2026: Gemini 3.6 Flash for multimodal reasoning and agents, 3.5 Flash-Lite for high-volume and low-latency work, and 3.5 Flash Cyber, a variant focused on vulnerability remediation.
The central strategic decision: price. Gemini 3.6 Flash costs $1.50/1M input tokens and $7.50/1M output tokens — below the $1.50/$9.00 of the previous 3.5 Flash version. Combined with greater token efficiency, the total cost per agentic task drops by up to 65% in long-horizon engineering tasks.
The numbers that matter
On the DeepSWE benchmark, 3.6 Flash rises from 37% (version 3.5) to 49% — a 12 percentage point increase. On MLE-Bench, it rises from 49.7% to 63.9%. On OSWorld-Verified — a computer use benchmark — it rises from 78.4% to 83.0%.
The 3.5 Flash-Lite, focused on volume, processes 350 output tokens per second — the fastest in the 3.5 series. Price: $0.30/1M input tokens and $2.50/1M output tokens.
The most relevant architectural change
Google integrated computer use directly into the Gemini API and Gemini Enterprise as a native client-side tool — eliminating the custom intermediary software that engineers previously had to build to allow models to operate on an operating system.
Both models have a 1 million token context window and a 64,000 token output limit. Knowledge cutoff: March 2026.
The elephant in the room
Gemini 3.5 Pro — the flagship model Google had signaled for summer 2026 — was not launched. Gemini 3.1 Pro, its predecessor, was launched in February 2026. Rivals OpenAI and Anthropic have launched several generations of more powerful flagship models since then. Google's Logan Kilpatrick responded on X: "Gemini 3.5 Pro is currently in testing with partners and we plan to make it broadly available once it's ready."

