Google DeepMind Unveils Gemini 4 Argon, Its Next-Generation Frontier AI Model
Key Takeaways
- •Gemini 4 Argon posts state-of-the-art benchmark results across domains, scoring 77.9% on DeepSWE v1.1 for long-horizon software engineering, 51.3% on Zapier's AutomationBench, and 91.7% on LVBench for long video understanding.
- •The model's output token limit has been expanded to an industry-leading 1 million tokens, roughly fifteen times the previous 64,000, which Google says allows difficult problems to be solved in a single pass.
- •Argon can autonomously discover, validate, and patch critical vulnerabilities, and via Wiz's Scan for Good initiative it exposed a critical flaw leaking sensitive personal information in healthcare software used by hospitals worldwide that earlier frontier models had missed.
- •Availability follows a phased rollout: trusted cyber defenders receive initial access through the Fairwind Program without cyber guardrails, while developers, enterprises, and consumers will gain access later, beginning with paid API customers and Google AI Ultra subscribers.
- •Inside Google, Argon agents beat a published quantum algorithm baseline by 40%, freed more than 300 TiB of data center memory, and are migrating large C/C++ codebases to Rust, including over 800,000 lines of the Fuchsia OS Zircon kernel.

Google DeepMind on September 30, 2026 announced Gemini 4 Argon, a new frontier AI model built to sustain deep reasoning across complex, long-horizon professional workflows. The company said the model delivers frontier performance in real-world software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense, and it is initially rolling out to a set of trusted cyber defenders through Google's Fairwind Program.
Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect at Google, announced the model in a blog post, writing that Argon is "fundamentally changing the way we work and build at Google." The company highlighted the model's industry-leading 1 million token limit for deep, multi-step problem solving, along with strengths in coding, financial research, legal drafting, and autonomous patching of cybersecurity vulnerabilities.
Phased Release and Pricing
Google said that safely releasing frontier capabilities at this level requires a phased approach. The company is actively engaged in the U.S. government's voluntary process for pre-release model access while it gradually expands availability, and it will continue gathering feedback from early testers as it iterates on guardrails before making Argon available to developers, enterprises, and consumers. The initial access plan means the capabilities and safeguards described by Google are being evaluated first in controlled settings rather than offered immediately to the broader market.
Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off the input token price. After the introductory period expires, pricing of $4 per 1 million input tokens and $20 per 1 million output tokens will apply. The broader rollout will begin with paid API customers and Google AI Ultra subscribers.
Changing How Google Works and Builds
Gemini 4 Argon is already powering internal workflows at Google, with thousands of Googlers highlighting the model's strengths in specialized coding tasks, deeper research, and writing quality. Google said the model is helping teams build faster, pushing the boundaries of engineering productivity and accelerating breakthroughs. Examples cited in the post include:
- Quantum algorithmic optimization: Argon is helping Google's quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
- Memory efficiency: A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google's data centers, freeing up more than 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
- Large-scale codebase migrations and optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries such as re2 and libgav1 up to more than 800,000 lines for the Fuchsia OS Zircon kernel. Given the criticality of many of these systems, the large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production. For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32,000 lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, and producing safe Rust so the compiler would vectorize it automatically. The result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++ implementation.
A 1 Million Token Output Limit
To support Gemini 4 Argon's capabilities across longer, more complex use cases, Google said it is expanding the model's output token limit to an industry-leading 1 million tokens, up from the previous 64,000 tokens. According to the company, when the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of reasoning depth that allows tough problems to be solved in one pass.
Coding and Enterprise Workflows Across Domains
Google said Gemini 4 Argon's capabilities across coding, reasoning, and multimodality, together with its ability to sustain long, multi-step tasks, allow it to excel across a range of enterprise workflows. Google engineers have been using the model for daily tasks ranging from everyday debugging to large-scale codebase migrations and algorithm design. Argon sets a new state of the art on DeepSWE v1.1, which measures a model's performance in real-world long-horizon software engineering tasks, with a score of 77.9%.
Beyond coding, Google said Argon is the leading model on the Vals Index, which measures economic impact across finance, coding, legal, and tax work, with every sector weighted by its contribution to U.S. GDP. The company reported similarly leading performance on domain-specific evaluations including Vals Finance Agent v2, which tests multi-step financial research, and Harvey's Legal Agent Benchmark, which covers legal research and drafting. On AutomationBench, Zapier's benchmark measuring end-to-end execution across core business functions, Argon ranks first with a score of 51.3%.
Argon is also strong when knowledge work requires visual understanding, according to Google. The model can drive professional chart analysis, identify details from long videos, and take action based on a series of documents. On LVBench, which measures long video understanding, the company said Argon is state of the art with a score of 91.7%.
Leading in Defensive Cybersecurity
Google said it trained Gemini Argon to be highly capable at cybersecurity defense to better equip cyber defenders for a new era of cyberattacks. The model can autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and Google's own internal teams, the company will release Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.
Security firm Wiz is already using Argon for cybersecurity defense through its Scan for Good initiative, a program dedicated to protecting critical public infrastructure for free by finding and remediating high-risk exposures. In an early demonstration of the model's impact, Google said, Argon uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, identifying a severe risk that previous frontier models had missed.
On CWE-bench v1, which evaluates a model's ability to remediate security vulnerabilities, Argon ties for first place with a top score of 68%, building on 3.8 Flash Cyber's frontier performance on CWE-bench v0.
Google said Gemini 4 Argon demonstrates impressive leaps in vulnerability discovery over 3.8 Flash Cyber:
- On Google's internal comprehensive vulnerability benchmark, Argon uncovered a wide range of exposures across complex codebases spanning 20 programming languages.
- On Wiz's internal black-box penetration testing benchmark, which tests a model's ability to analyze live web systems without source code, Argon outperforms 3.8 Flash Cyber in discovering the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence to validate them.
Strengthening Frontier Safeguards
Before rolling out Gemini 4 Argon broadly, Google said it is continuing to strengthen critical frontier safeguards across four main areas:
- Defending against misuse: To prevent bad actors from using Argon for cyber or chemical, biological, radiological, and nuclear (CBRN) attacks, the model is designed to refuse harmful requests while preserving legitimate, dual-use scientific research, as per Google DeepMind's Frontier Safety Framework. The company is strengthening the robustness of its safeguards for this launch, including improving techniques to monitor the model's internal activations to spot misuse. The safeguards underwent robustness testing by internal and external red teams using a combination of manual and automated attack methods.
- Defending against prompt injection attacks: Google described Argon as its most resilient model yet against indirect prompt injections, in which malicious instructions or context are used to hijack a model's behavior. The company said these are complex attacks that require constant vigilance and multiple layers of defense. Through automated red teaming and adversarial training, Gemini 4 Argon leads in prompt injection robustness on Gray Swan's Indirect Prompt Injection (IPI) benchmark, according to Google.
- Monitoring for misalignment: To prevent Argon from stepping out of bounds while accomplishing a task in ways that go beyond user intentions, Google is deploying misalignment mitigations that monitor the model's chain-of-thought and actions and stop execution when necessary. The company used a similar system to monitor its training runs and send alerts to a dedicated incident response team, taking careful precautions against feeding the findings back into training so as not to risk shaping Argon's reasoning to evade monitoring. Google said it strongly encourages the rest of the industry to preserve reasoning transparency "in these pivotal moments of increased capabilities while navigating alignment risks," so that model thoughts remain helpful in identifying and diagnosing misalignment.
- Hardening systems: As frontier models grow increasingly capable, Google said safely testing them requires secure environments that can keep up with the systems themselves. In line with its agent control roadmap, the company is hardening its sandboxed environments by isolating and sealing them before high-risk training or evaluations begin. Google also committed to sharing these agent security best practices with partners to improve security across the ecosystem.
Rolling Out Soon
Google said it built Gemini 4 Argon with frontier-level capabilities in coding, knowledge work, cybersecurity defense, and creative writing to serve as a partner for developers, professionals, and enterprises tackling their most difficult problems. The company expressed gratitude to the initial cohort of cyber defenders and trusted testers whose real-world evaluations and feedback will help strengthen its systems before release to developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers.
About the author: Koray Kavukcuoglu is one of the world's foremost experts in artificial intelligence. As SVP of Google DeepMind, he leads the development of Google DeepMind's state-of-the-art generative AI models and their integration into Google's products at scale. Prior to his current role, he was VP of Research at Google DeepMind, where he founded the deep learning team, which pioneered major research breakthroughs such as DQN (Deep Q-Network) and WaveNet.