Google DeepMind shipped Gemini 3.8 Flash on September 2, alongside a sibling called Gemini 3.8 Flash Cyber. Same foundational model underneath both. The difference is not parameter count, context length, or training data. It’s which safety mitigations are turned on, and who gets a key.
That’s a genuinely new shape for a model release, and it deserves more attention than the version number is getting.
Two doors, one model
Regular 3.8 Flash is generally available at $0.75 per million input tokens and $3.75 per million output, introductory pricing that runs through December 31, 2026. Flash Cyber is gated to vetted defenders and posts 47.2% pass@1 on CWE-Bench, a vulnerability-focused benchmark.
Until now, the industry’s answer to “this model is too capable for open release” was to not release it, or to release a lobotomized version and hope nobody noticed. Google’s answer here is a tiered access envelope: same intelligence, different guardrails, different application process. Anthropic and OpenAI have both flirted with this. Google just made it a product SKU.
The honest read is that this is a good idea with a soft center. Vetting is the entire security model. If the process for becoming a “vetted defender” is a form and a corporate email domain, the fence is decorative. Google hasn’t said much about what vetting involves, and that silence is the part worth pressing on. A 47.2% pass rate on finding vulnerabilities cuts both ways depending on who’s holding it.
The other half is a pricing story
Meanwhile, the mainline 3.8 Flash arrived only a few weeks after 3.7 Flash. Google’s pitch, per The Verge, is that it “works harder”: more reasoning steps on complex tasks, calling tools iteratively rather than answering in one shot.
Same sticker price as 3.7. Different bill.
If the model burns more tokens per request by design, the per-million rate is a number that no longer maps to what you actually pay. A task that cost you 2,000 output tokens on 3.7 might cost 6,000 on 3.8, and your invoice triples while the price sheet stays frozen. Google is not hiding this; “works harder” is the marketing line. But calling it “same introductory pricing” while the token consumption profile shifts underneath is a sleight of hand that every lab is now performing, and buyers keep falling for it.
The fix is boring. Benchmark 3.7 versus 3.8 on your own workload, measure total cost per completed task, and ignore the rate card entirely. If 3.8 solves problems 3.7 couldn’t, paying more is fine. If it just thinks longer to reach the same place, you’ve been upsold.
Two questions I’d want answered before December 31, when that introductory pricing expires: what does Google’s defender vetting actually check, and how many tokens does 3.8 Flash spend on a task 3.7 Flash already handled.