NewsStocksGoogle DeepMind Unveils Gemini 3.8 Flash and 3.8 Flash Cyber for Agentic Workflows and Cybersecurity

Google DeepMind Unveils Gemini 3.8 Flash and 3.8 Flash Cyber for Agentic Workflows and Cybersecurity

Author: Google DeepMind Blog·

Key Takeaways

  • Google DeepMind released Gemini 3.8, its most capable reasoning and coding model, three weeks after 3.7 Flash, marking its third Flash release in six weeks.
  • Gemini 3.8 Flash is offered at an introductory price of $0.75 per million input tokens and $3.75 per million output tokens, with rates doubling on January 1, 2027.
  • Gemini 3.8 Flash Cyber achieved frontier-level vulnerability discovery on CyberGym and over 70% success on Google's internal 20-language vulnerability benchmark.
  • The Chrome Security team found Flash Cyber produced 2.6 times more correct patches to Chrome vulnerabilities than larger commercial models.
  • Flash Cyber is restricted to trusted defenders through the new Fairwind Program, while 3.8 Flash is available via the Gemini API, Gemini Enterprise, and Google AI Pro and Ultra subscriptions.
Google DeepMind Unveils Gemini 3.8 Flash and 3.8 Flash Cyber for Agentic Workflows and Cybersecurity

Google DeepMind has introduced Gemini 3.8, its most capable reasoning and coding model to date, arriving just three weeks after 3.7 Flash and marking the company's third Flash release in six weeks. The accelerated cadence reflects the current pace of the frontier-model market, where Google, OpenAI, and Anthropic have been shipping successive releases at intervals of weeks rather than quarters, and where mid-tier "Flash"-class models are increasingly the workhorses of production AI applications. The launch includes two variants: Gemini 3.8 Flash, a general-purpose workhorse model, and Gemini 3.8 Flash Cyber, a specialized cybersecurity model available to vetted defenders through the new Fairwind Program.

According to the announcement by Tulsee Doshi, Senior Director of Product Management, and Raluca Ada Popa, Gemini Security Lead at Google DeepMind, the two releases are powered by the same foundational intelligence, accelerated by long-running agentic loops that recursively evaluate and refine the underlying models. The coding and reasoning gains across this shared core were driven by several innovations, including rigorous training in the demanding domain of cybersecurity. The pairing is notable because it applies a technique — using a hard, verifiable domain like security as a training signal — that has also been used elsewhere in the industry to sharpen general model reasoning.

Gemini 3.8 Flash: built for long-horizon coding and autonomous agents

Gemini 3.8 Flash delivers substantial improvements over 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning in specialized domains, often approaching the performance of higher-cost frontier models. It is available at the same introductory price as 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens. This introductory pricing expires on December 31, 2026; from January 1, 2027, rates of $1.50 per million input tokens and $7.50 per million output tokens will apply. Even at the full rate, the model is priced well below typical frontier-model tiers, underscoring the competitive pressure on per-token pricing as vendors compete for agentic workloads that consume far more tokens than chat-style interactions.

On DeepSWE v1.1 (Long-Horizon Software Engineering), 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, at a fraction of the cost. The model also demonstrates dependability for critical enterprise autonomy across specialized knowledge domains. In quantitative and professional fields requiring advanced analysis and reporting, 3.8 Flash outperforms 3.7 Flash and other frontier models on benchmarks such as Vals Finance Agent V2 and Harvey's Legal Agent Benchmark. It also achieves 54.9% on HLE-Verified, reflecting multi-step reasoning capability across STEM, humanities, and professional fields.

These gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, the model executes additional reasoning steps and calls tools iteratively, sometimes consuming more tokens to maximize performance, particularly at higher effort levels. This trade-off between accuracy and token spend has become a central design question for agentic AI, since autonomous agents that loop over many steps can multiply inference costs; the effort-level control gives developers a lever over that trade-off. For compute-constrained applications, developers can use lower effort levels to reduce token overhead or continue relying on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads.

Google highlighted several demonstrations built with 3.8 Flash in Google Antigravity: a 3D castle-navigating wizard game created from a simple looping prompt, with puzzles, environmental storytelling, and textures generated by Nano Banana; a fully playable DOS version of Google Maps with locations, directions, and Street View; a topographic map of famous geographical sites featuring real-time cross-sections, 2D projections, and scientific explanations using real U.S. Geological Survey datasets; and Hardware Anatomy, an interactive 3D visualizer built in Google AI Studio that generates realistic Three.js renderings of physically proportioned hardware teardowns with an explodable-layer deconstruction slider.

Gemini 3.8 Flash Cyber: expert cyber performance

Gemini 3.8 Flash Cyber provides defenders with frontier-level capabilities in vulnerability detection and automated patching, delivered at Flash-tier speed and cost to enable rapid iteration. It arrives amid a broader industry push toward AI-assisted security, as software supply-chain attacks and unpatched vulnerabilities at large organizations have made automated discovery and remediation an active area of investment across the security industry — while also raising concerns that similar capabilities, if unrestricted, could aid attackers.

Autonomous vulnerability discovery. On CyberGym, the standard industry benchmark for finding vulnerabilities, 3.8 Flash Cyber demonstrates frontier-level autonomous vulnerability discovery, surpassing both 3.5 Flash Cyber and significantly larger frontier models. To better reflect real-world defensive needs beyond the C/C++ codebases covered by CyberGym, Google also evaluated the model on a comprehensive internal benchmark requiring vulnerability discovery across complex codebases spanning 20 programming languages, where it achieved a success rate exceeding 70% — a substantial leap over previous models.

Automated patching. Google said it focused on equipping defenders with expert capabilities that favor them over attackers, investing in vulnerability fixing from the start and prioritizing it over offensive capabilities such as exploitation. On CWE-Bench, a challenging external patching benchmark run by Collinear, 3.8 Flash Cyber sits on the Pareto frontier with a pass@1 of 47.2% versus 47.8% for a leading frontier model — at significantly lower cost.

Real-world impact. Google reported it is already using 3.8 Flash Cyber to secure code across the company. The Chrome Security team found the model produced 2.6 times more correct patches to Chrome vulnerabilities than much larger best-in-class commercial models. Wiz found it achieved 7.5–9.7% higher recall on its internal penetration testing benchmark at 2.3–5.2x lower cost compared with other leading frontier models. Google's Cloud Vulnerability Research team used the model to find a critical foundational vulnerability in under two hours — research and discovery for which the process typically takes months. Third-party validation from Wiz and Collinear's external benchmark provides some independent corroboration of Google's claims beyond self-reported internal evaluations, though detailed benchmark methodologies remain vendor-published.

Built with safety in mind

Gemini 3.8 Flash ships with safeguards against misuse in Chemical, Biological, Radiological, and Nuclear (CBRN) domains and cyber offense, while permitting beneficial use cases, in line with Google's Frontier Safety Framework. 3.8 Flash Cyber carries a more permissive set of cybersecurity mitigations and is accordingly restricted to trusted defenders who require a more comprehensive set of cyber capabilities. This restricted-access approach mirrors a wider debate in the AI and security communities over how to distribute powerful offensive-capable cyber capabilities — limiting them to vetted defenders — without stifling legitimate security research. Google also noted that the Gemini 3.8 models have made a significant leap in prompt injection robustness as measured by Gray Swan, protecting users from prompt-injection-related malicious attacks; prompt injection remains one of the most-cited security risks for agentic AI systems that read untrusted content and take actions on a user's behalf.

Availability

  • Developers: Build with 3.8 Flash and explore agent-first workflows in Google Antigravity, or start building in the Gemini API via Google AI Studio and Android Studio, or generate UIs in Stitch. Developer documentation is available via Google's developer docs.
  • Enterprises: Access 3.8 Flash in Gemini Enterprise.
  • Consumers: 3.8 Flash is available to Google AI Pro and Ultra subscribers across the Gemini app, AI Mode in Google Search, and Gemini in Google Sheets.
  • Cyber: Through the new Fairwind Program, trusted government authorities, critical infrastructure operators, and software maintainers receive prioritized access to Gemini 3.8 Flash Cyber. Organizations can apply for access.

Source: Google DeepMind Blog