4Geeks chosen to deliver AI education in the Bahamas alongside Harvard, Oxford, and Columbia.See more
9 min read

Gemini 4 Argon: Google's Model That Finds and Fixes Flaws

Gemini 4 Argon is Google's new frontier model. What it does, how it finds and patches vulnerabilities, who can use it today and what it will cost.

Short answer: Gemini 4 Argon is Google's new frontier model, announced on September 30, 2026, for long and complex work in software engineering, legal and finance research, and cybersecurity defense. Its headline skill is security: Google says Argon can find, validate and patch critical software vulnerabilities on its own. Because the same skills could help attackers, Google is releasing it first to vetted cyber defenders through its Fairwind Program, with paid API customers and Google AI Ultra subscribers next (Google).

Gemini 4 Argon finding and sealing a flaw in a codebase

Most frontier models open to everyone on launch day. Argon doesn't. Google is treating its first Gemini 4 model as a tool that can harden software or break it, depending on who holds it, and it decided who goes first. This guide covers what Argon does, how its vulnerability work is described, why access is restricted, when you can expect to use it and what it will cost. For the wider landscape of assistants and agents, start with our AI tools hub.

What is Gemini 4 Argon?

Gemini 4 Argon is the first model of Google's Gemini 4 generation. Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect, introduced it as a model built for deep reasoning across long-horizon workflows: tasks that take many steps and a lot of context, such as migrating a large codebase or producing a full legal or financial analysis.

The most visible technical change is output length. Argon can write up to 1 million tokens in a single response, up from the 64,000-token limit of earlier Gemini models. In practice, that lets the model deliver a large piece of work in one pass instead of splitting it into dozens of smaller requests that you have to stitch back together.

Google backs the claim with internal examples. According to the announcement, Argon helped migrate more than 800,000 lines of code for the Zircon kernel in Fuchsia and made a video decoder 2.7 times faster.

What can Argon do with code and long tasks?

Google published scores across coding, automation and enterprise work. These are Google's own reported results, not independent tests:

BenchmarkWhat it measuresGemini 4 Argon
DeepSWE v1.1Long, real-world software engineering tasks77.9%
CWE-bench v1Fixing software vulnerabilities68%, tied for first
AutomationBenchAutomated multi-step workflows51.3%, ranked first
LVBenchUnderstanding long videos91.7%

On CWE-bench v1, Argon shares first place with OpenAI's GPT-6 Astra and xAI's Grok 4.7, according to SecurityWeek. Google also says Argon leads Gray Swan's indirect prompt injection benchmark, which tests how well a model resists instructions hidden in the content it reads. That matters for any agent that browses the web or reads tickets and email.

On the enterprise side, Google reports leading results on the Vals Index, the Vals Finance Agent benchmark and Harvey's Legal Agent Benchmark.

How does Argon find and patch vulnerabilities?

Google describes Argon's security work in three verbs: find, validate and patch.

  • Find: read a codebase and spot code that could be exploited, from a missing input check to a flaw in how memory is handled.
  • Validate: confirm that the flaw is real and reachable, so the team isn't buried in false alarms. A finding you can reproduce is worth far more than a long list of maybes.
  • Patch: propose a code change that closes the hole without breaking what the code is supposed to do.

Google tested this on an internal benchmark of complex codebases written in 20 programming languages. The most concrete example came from Wiz, a cloud security company, through its Scan for Good initiative: Argon found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals around the world, a flaw that previous frontier models had missed. The vendor and the patch status were not disclosed (Help Net Security). On Wiz's black-box penetration testing benchmark, Argon also beat Gemini 3.8 Flash Cyber at mapping what an attacker could reach and at producing proofs of concept.

Argon builds on a model Google released a few weeks earlier. In September, Gemini 3.8 Flash Cyber went to vetted defenders, and Google's Chrome Security team reported that it produced correct vulnerability patches at 2.6 times the rate of the best comparable commercial models (Tech Times).

Why did Google restrict access to Argon?

Because security skills are dual-use. A model that can find a flaw, prove it works and write the fix can also hand an attacker a working exploit. Google chose to give defenders a head start:

  • Fairwind Program first. Google launched Fairwind on September 3, 2026, alongside Gemini 3.8 Flash Cyber, for government agencies, critical infrastructure operators, maintainers of widely used software and vetted security partners. Argon is rolling out to these trusted defenders now.
  • A version without cyber guardrails. Fairwind members and Google's own internal teams get Argon without its cyber restrictions, so they can use its full defensive capabilities.
  • Government review. Google says it is taking part in the U.S. government's voluntary process for pre-release model access.

The public version comes with safeguards. Google says Argon is designed to refuse requests that would enable cyber, chemical, biological, radiological or nuclear attacks while still supporting legitimate dual-use research. Google's announcement also lists monitoring of the model's internal activations, automated red teaming against prompt injection, monitoring of its chain of thought for misaligned behavior and hardened sandboxes for agents.

OpenAI took a different route with GPT-6 Astra, released on September 3. Astra reached ChatGPT Plus, Pro, Business and Enterprise users and the API within days, but it refuses more advanced cybersecurity tasks, such as writing proof-of-concept exploits. OpenAI says it plans to widen access with less restrictive safeguards for defensive work through its Daybreak program (OpenAI). Put simply, OpenAI released the model and held back the risky capabilities, while Google held back the model itself.

Who can use Gemini 4 Argon today, and when will everyone get it?

StageWho gets accessWhat they get
NowTrusted defenders in the Fairwind Program and Google's internal teamsArgon without cyber guardrails
NowU.S. government voluntary pre-release reviewAccess to evaluate the model before wide release
NextPaid Gemini API customers and Google AI Ultra subscribersArgon with its standard safeguards
LaterDevelopers, enterprises and consumers more broadlyNo date yet. Google says "as soon as possible"

If you are not a security defender, you can't use Argon yet. Google says it will expand access gradually based on what it learns from early testers.

How much will Gemini 4 Argon cost?

For developers, Google has published API prices. Argon launches at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens 95% cheaper than regular input. After the introductory period, the price rises to $4 per million input tokens and $20 per million output tokens.

For consumers, the first route will be the Google AI Ultra subscription. Google hasn't said when Argon will reach the other Gemini plans.

Keep the 1 million token output limit in mind when you estimate costs. A single long answer can consume a lot of output tokens, so test on real tasks before you plan a budget.

What does Argon mean for developers and security teams?

Models like Argon and Astra change how software gets secured. A few practical consequences:

  • More AI-found bug reports. Maintainers of popular projects should expect more vulnerability reports produced with AI. Validated findings help, but triage still takes people and time.
  • AI patches still need review. A patch that closes one hole can open another or quietly change behavior. Run tests, check the diff and keep a human in the approval loop. Our guide on how to review and verify AI-generated code covers the checks that matter.
  • Basics matter more, not less. When models find the easy flaws in minutes, clean input validation, up-to-date dependencies and least privilege stop being optional.
  • Agents need limits. A model with access to your code and tools should run in a sandbox, with scoped permissions and logs. The OpenBot governance case shows what that isolation looks like in practice.

How can you prepare before Argon opens up?

  1. Check whether Fairwind applies to you. If your organization is a government agency, a critical infrastructure operator or the maintainer of widely used software, Fairwind is the route Google has opened first.
  2. Build a small test set. Collect a few real bugs and vulnerabilities your team has already fixed. When Argon reaches the API, you can see whether it finds them and how good its patches are.
  3. Set up review gates now. Automated tests, static analysis and a human approval step should sit between any AI patch and your main branch.
  4. Compare what you can use today. While you wait, today's coding agents already handle a lot of this work. Our comparison of the best AI coding agents shows which tools fit which jobs.

The skill behind using a model like Argon

Argon doesn't remove the need for people who understand code and security. It rewards them. Someone still has to decide what to scan, judge whether a finding is real and approve a patch before it ships. Those are skills you can learn. If you want to build them with mentors and real projects, compare our AI programs and find the one that fits your starting point.

Learn to build with models like Argon

Tell us what you code today and an advisor will help you see which program fits: AI Engineering, AI Engineering for Developers or AI Flex.

Frequently Asked Questions