Google just shipped three new Gemini models in a single drop: Gemini 3.6 Flash, 3.5 Flash-Lite, and a gated 3.5 Flash Cyber. Notably absent is the one everyone actually wants to talk about — Gemini 3.5 Pro, which remains delayed. The release tells two stories at once: a genuinely aggressive push on the economics of mid-tier inference, and a flagship strategy that looks increasingly cautious.

The economics are the headline

Strip away the version-number soup and the interesting move is cost. Gemini 3.6 Flash cuts output tokens by roughly 17% and drops its output price to $7.50 per 1M tokens. That combination — fewer tokens and a lower per-token rate — compounds. For anyone running agentic workloads, where a single task can chain dozens of model calls, token efficiency isn’t a rounding error; it’s the difference between a workflow that pencils out and one that doesn’t.

Flash-Lite leans the other direction, chasing throughput at 350 tokens/sec. That’s a latency play, aimed at the interactive and high-volume cases where speed beats raw reasoning. Together the two models bracket the mid-tier: one optimized for cheap agent loops, one for fast responses. It’s a coherent segmentation, and it reads as Google acknowledging that most real production traffic never touches the flagship at all.

Flash Cyber is the quiet standout

The gated 3.5 Flash Cyber model is the most intriguing entry. It powers CodeMender, Google’s vulnerability-finding system, and its restricted availability signals Google treating offensive-security-adjacent capability as something to meter rather than broadcast. A Flash-class model tuned for security work is a bet that agentic code analysis is now good enough — and cheap enough — to run at scale across large codebases. That’s a more concrete near-term application than most flagship demos.

The missing Pro is the story

Here’s where skepticism is warranted. Shipping three mid-tier models while the flagship slips raises fair questions about Google’s roadmap. Are they optimizing the profitable, high-volume part of the stack because that’s where the money is? Or is 3.5 Pro genuinely hard to land — held back by cost, safety review, or benchmarks that aren’t beating the competition convincingly enough to ship?

Either reading is plausible, and Google isn’t saying. What’s clear is that the company is comfortable letting the Flash tier carry the release narrative. That’s a defensible commercial choice — the mid-tier is where developers live — but it leaves a strategic vacuum at the top of the lineup that rivals will happily fill.

What to actually do

If you’re building agents, 3.6 Flash’s efficiency gains are worth benchmarking against your current stack today; the cost curve matters more than the version bump. If you were holding out for 3.5 Pro, keep holding — and keep your options open. A flagship that ships late and quiet is one you should evaluate on results, not anticipation.