Google says all three of its new Gemini models are the most advanced it’s ever shipped. It says that every time. So I went looking for the one number that actually matters, and I found it buried in a token-usage stat most people will scroll right past.
![Google dropped 3 new Gemini models and only one earns the hype [2026] Google dropped 3 new Gemini models and only one earns the hype [2026]](https://cryptogames.gg/wp-content/uploads/2026/07/google-dropped-3-new-gemini-models-and-only-one-earns-the-hype-2026-1.png)
The lineup is Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber. Three models, three jobs, and only one of them changes what you’ll pay to run it at scale.
The 17% that pays your bill
Start with 3.6 Flash, the one Google calls its workhorse. The pitch is better coding, better knowledge work and better multimodal handling than the model it replaces.
Here’s the part worth your attention. According to the Artificial Analysis Index, 3.6 Flash cuts output token usage by 17% compared to 3.5 Flash. In some benchmarks, like DeepSWE by Datacurve, Google says it sees up to 65%.
![Google dropped 3 new Gemini models and only one earns the hype [2026] Google dropped 3 new Gemini models and only one earns the hype [2026]](https://cryptogames.gg/wp-content/uploads/2026/07/google-dropped-3-new-gemini-models-and-only-one-earns-the-hype-2026-2.png)
If you’ve ever watched a Flash model burn tokens narrating its own reasoning before answering, you know why this lands. Fewer output tokens at a lower cost per output token isn’t a headline feature. It’s the line item on your invoice.
In Google’s own words: “3.6 Flash: Our workhorse model that delivers better coding, knowledge work, and multimodal performance. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash, and in some benchmarks like DeepSWE by Datacurve, we observe up to 65%, all at a lower cost per output token.”
![Google dropped 3 new Gemini models and only one earns the hype [2026] Google dropped 3 new Gemini models and only one earns the hype [2026]](https://cryptogames.gg/wp-content/uploads/2026/07/google-dropped-3-new-gemini-models-and-only-one-earns-the-hype-2026-3.png)
Flash-Lite is built for one thing: speed
Then there’s 3.5 Flash-Lite, aimed at low-latency work, chat and document processing. Google calls it the fastest, most cost-effective model in its 3.5 class.
The number here is 350 output tokens per second, again per the Artificial Analysis Index. Google also says it beats prior Flash-Lite generations in agentic workflows, and specifically that it “significantly” outperforms its predecessor at the thinking level.
That’s the model you reach for when a user is waiting on a reply and every second of latency costs you. It’s not trying to out-reason the workhorse. It’s trying to answer before you notice the delay.
Google’s line: “3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model, delivering 350 output tokens per second according to the Artificial Analysis Index, also significantly outperforming prior Flash-Lite generations in agentic workflows.”
The security one is a model plus an agent, not just a model
The third, 3.5 Flash Cyber, is the one people will misread. It’s fine-tuned to help find and fix cybersecurity vulnerabilities, at a lower cost per token than the larger models.
But Google is upfront that the model alone isn’t the story. It pairs a specialized cyber-focused model with its CodeMender code security agent, and it’s clear about why: security work needs the model orchestrated alongside agent infrastructure, not dropped in on its own.
Google’s description: “3.5 Flash Cyber in CodeMender: Successful cybersecurity applications require careful orchestration of a model alongside an agent infrastructure. We’re introducing a combination of a new, highly efficient, specialized cyber-focused model paired with our CodeMender code security agent that delivers competitive performance at the frontier.”
Note the phrasing. “Competitive performance at the frontier,” not best in the world. When a company hedges its own claim, believe the hedge.
Which one to actually use
If you’re running a coding assistant or anything multimodal at volume, 3.6 Flash is the one to test first, and the 17% token cut is the reason. If your app lives and dies on response time, Flash-Lite and its 350 tokens per second is the pick. And if you’re doing vulnerability work, don’t evaluate Flash Cyber in isolation, because Google didn’t design it to run that way.
Three models, one that moves the cost math and two that stay in their lane. Google says you can learn more on its announcement page.
Source: GSMarena.