Google has launched Gemini 3.8 Flash alongside a security-focused sibling, Gemini 3.8 Flash Cyber, the company announced on September 2, 2026. It is the third Flash-family release in just six weeks, arriving three weeks after 3.7 Flash. Google's framing is unusually direct: "our best reasoning & coding model yet, at the same speed and low cost of 3.7" — a capability jump without a price jump, at least until the calendar flips to 2027.

What exactly was announced?
The post was authored by Tulsee Doshi, Senior Director of Product Management at Google, and Raluca Ada Popa, Gemini Security Lead at Google DeepMind. Two distinct models carry the 3.8 badge, aimed at very different audiences:
Gemini 3.8 Flash: Google calls it "our most intelligent workhorse model," with significant gains over 3.7 Flash in software engineering, agentic tasks, and multi-step reasoning across specialized domains. The introductory API price matches the previous generation at $0.75 per million input tokens and $3.75 per million output tokens.
Gemini 3.8 Flash Cyber: A cybersecurity specialist that detects vulnerabilities and patches them automatically. It is available only to "trusted defenders" — governments, critical infrastructure operators, and software maintainers — through the new Fairwind Program. Content creators and independent developers will not touch this one.
The numbers behind the claim
The benchmarks Google chose to highlight say a lot about where users will actually feel the difference — long, multi-step, specialized work:
- DeepSWE v1.1: On long-horizon software engineering, 3.8 Flash outperforms most larger, more expensive frontier models at solving complex engineering problems end to end, at a fraction of the cost.
- HLE-Verified at 54.9%: a multi-step reasoning benchmark spanning STEM, humanities, and professional fields.
- Professional domains: it leads 3.7 Flash and rival frontier models on the Vals Finance Agent V2 benchmark and on Harvey's Legal Agent Benchmark — exactly the kind of analytical, report-shaped work that gets delegated to AI assistants in real workplaces.


A model that "works harder" — and what that costs you
The most candid paragraph in the announcement explains the performance jump: 3.8 Flash "works harder." On complex tasks it executes extra reasoning steps and calls tools iteratively, sometimes spending more tokens to maximize performance — especially at higher effort levels. Developers who need compute efficiency can dial effort down or stay on 3.7 Flash, which Google says "remains fully supported for efficiency-first workloads." In plain terms: the extra intelligence is not free at the consumption level, even though the per-token price is unchanged.
Where you can try it today
For non-developers, this is the part that matters, in Google's own words: 3.8 Flash is available to Google AI Pro and Ultra subscribers across the Gemini app, AI Mode in Google Search, and Gemini in Sheets. Developers can build with it through Google AI Studio, Antigravity, and Stitch, while enterprises get it inside Gemini Enterprise. The announcement demos lean into ambition: a fully playable 3D wizard-castle game with textures generated by the Nano Banana image model, a functional DOS-style version of Google Maps, and an interactive USGS topographic map — all produced from single prompts.
What this means for content creators
- Your Gemini assistant got smarter overnight: if you already pay for AI Pro or Ultra, the heavy tasks — digesting long research, structuring a content strategy, analyzing data — now get deeper processing without changing how you work.
- Your favorite tools may quietly improve: any app built on the Gemini API inherits near-frontier performance at mid-tier pricing, which tends to raise the baseline across the whole writing and analytics tool market.
- Budget carefully if you build: "working harder" can mean more tokens per task at high effort levels. Watch your bill before committing a product to this model.
- Mark the calendar: the introductory price expires December 31, 2026. From January 1, 2027 it doubles to $1.50 per million input and $7.50 per million output tokens. Model your unit economics on the new price, not the teaser.
Want to put this class of capability to work on long-form content? Try the ARWriter platform, which pairs frontier models with a full Arabic-first writing and publishing workflow, or read our deep dive on Qwen3.8-Flash, the open rival that undercuts everyone on price.
Quick comparison with what you use today
| Model | Price per 1M tokens (in/out) | Key point |
|---|---|---|
| Gemini 3.8 Flash | $0.75 / $3.75 (until end of 2026) | Best reasoning and coding at 3.7's price |
| Gemini 3.7 Flash | $0.75 / $3.75 | Still fully supported, the efficiency-first option |
| Qwen3.8-Flash (Alibaba) | $0.15 / $0.47 | Dramatically cheaper with a million-token context |
| Claude Fable 5.1 (Anthropic) | $10 / $50 | Premium tier at a steep premium |
The picture that emerges: Google is attaching near-frontier performance to mid-tier pricing, the cheaper Chinese option keeps pressure on cost-sensitive projects, and Anthropic holds the premium bracket. For a related shift in Google's consumer AI economics, see our explainer on Gemini Notebook's new flexible usage limits.
Safety notes worth knowing
Google says 3.8 Flash ships with safeguards against misuse in high-risk domains — chemical, biological, radiological, and nuclear — under its Frontier Safety Framework, and reports a significant leap in prompt-injection robustness as measured by Gray Swan. That last point matters to anyone building AI agents that touch untrusted content, since prompt injection remains the attack path of choice against automated workflows.
Honest limitations
- This launch is developer- and agent-first: the showcased benchmarks are coding and analytics. Nothing in the announcement promises new creative-writing or image-generation features — you get a smarter general model inside the same interfaces.
- The introductory price is temporary: the January 2027 doubling could reshuffle the value comparison, especially since 3.7 Flash stays available at its current price as the economy option.
- Higher effort means higher consumption: the per-token rate did not rise, but token counts can, and the two multiply each other on heavy workloads.
- Cyber is not for you: every advanced security capability sits behind the Fairwind Program's vetting wall.
Frequently asked questions
Is Gemini 3.8 Flash free to use?
The announcement scopes consumer availability to paid Google AI Pro and Ultra subscriptions across the Gemini app, AI Mode in Search, and Sheets. Developers access it via the Gemini API in AI Studio at the introductory price. No free-tier availability was mentioned.
What is the difference between Gemini 3.8 Flash and Flash Cyber?
3.8 Flash is the general-purpose model for everyone. Flash Cyber is a vulnerability-discovery and automated-patching specialist restricted to vetted defenders in the Fairwind Program — governments and critical infrastructure operators. It is simply not sold to the public.
Should I switch from 3.7 Flash to 3.8?
If your workload involves complex, multi-step reasoning — long analyses, code, professional reports — the announced gains are substantial. For short-form everyday writing, 3.7 Flash remains fully supported and lighter on tokens, which may make it the better default.
When does the introductory price end?
Per the official footnote, it expires December 31, 2026. From January 1, 2027, pricing becomes $1.50 per million input tokens and $7.50 per million output tokens — double today's rate.
Does 3.8 Flash improve other languages, such as Arabic?
The announcement makes no language-specific claims and lists no Arabic-language results; its evidence base is English-language reasoning and coding benchmarks. Gemini supports Arabic conversationally, but anyone with Arabic production workloads should run their own side-by-side tests before standardizing on it.
Bottom line
The 2026 model war is no longer just about who is smartest — it is about who delivers frontier intelligence at mid-market prices. For creators and developers, Gemini 3.8 Flash means heavier tasks get deeper treatment without changing tools or subscriptions, while anyone building on the API should read the pricing fine print twice before making long-term commitments.
Sources
- Official announcement on the Google blog — primary source (September 2, 2026)
- The Verge AI coverage — secondary confirmation