Tech • AI • Robotics • Game

VIDEO
ENFR

Google Just Dropped Argon: Their Most Powerful AI Ever

9.4/10
AIAI RevolutionOctober 2, 2026 at 01:41 AM13:39
Audio player
0:00 / 0:00

TL;DR

Google has unveiled Gemini 4 Argon, a new flagship AI model focused on software engineering, enterprise work, and cyber defense, with early access limited to trusted security teams because of its ability to find and fix serious vulnerabilities autonomously.

KEY POINTS

Phased launch centered on cyber defense

Gemini 4 Argon is rolling out first through Fairwind, Google’s security initiative for a small group of trusted cyber defenders. The company said the model’s capabilities require a phased release, with participation in the US government’s voluntary pre-release review process before broader availability. Access is set to expand later to paid API customers, Google AI Ultra subscribers, enterprises, developers, and consumers, though no date was given.

Pricing signals an aggressive opening

Google set an introductory price of $2 per million input tokens and $10 per million output tokens. Cached input tokens, for reused context, are discounted by 95%, or about $0.10 per million. The use of “introductory” suggests pricing may rise after launch.

Unrestricted version reserved for trusted defenders

Google said Argon was trained to handle cyber defense tasks end to end, including finding critical vulnerabilities, validating them, and patching them autonomously. Because the same knowledge used to fix flaws can also be used to exploit them, a guardrail-free version is being provided only to trusted defenders and Google’s internal teams. Standard versions are designed to refuse assistance for cyberattacks and other high-risk misuse.

Early real-world security result

In an early demonstration with cloud security company Wiz, Argon identified a critical flaw in healthcare software used by hospitals worldwide that exposed sensitive personal data. Google said earlier frontier models had missed the issue. The result is being presented as evidence that the new model can detect vulnerabilities beyond prior systems.

Benchmark gains in vulnerability remediation

On CWE Bench V1, which measures how well models remediate software vulnerabilities, Argon scored 68%, tying for first place against systems including GPT-6 Astra, Claude Opus 5.5, and Grok 4.7. Google also said Argon outperformed its earlier security model, Gemini 3.8 Flash Cyber, on internal tests and on Wiz black-box penetration benchmarks involving live web targets without source-code access.

Heavy internal use at Google

Google said thousands of employees have already been using Argon for coding, research, and writing. The company described applications ranging from debugging and large code migrations to algorithm design. In quantum computing work, Argon reportedly improved a published baseline by 40% in minutes by optimizing subroutines that reduce “space-time resources,” a measure combining qubits and gate operations.

Large code migrations and infrastructure savings

Argon agents are being used to migrate code from C and C++ to Rust, a language designed to prevent classes of memory-safety bugs. Google said projects range from libraries such as RE2 and libgav1 to the Fuchsia Zircon kernel, which has more than 800,000 lines of code. In another internal effort, Argon-driven memory optimizations across Google’s data center fleet freed more than 300 tebibytes of memory, with total savings estimated between 500 tebibytes and 1 pebibyte.

A notable Rust performance case

In work on libgav1, Google’s open-source video decoder, Argon agents replaced 32,000 lines of hand-tuned SIMD code with safe Rust structured for compiler vectorization. Google said the result was a memory-safe decoder that ran 2.7 times faster than the prior Rust port while preserving identical video output and narrowing the gap with the optimized C++ version.

Longer output and broad benchmark claims

Google raised Argon’s output limit to 1 million tokens, up from 64,000, to support long, multi-step workflows. On DeepSuite V1.1, a benchmark for long-horizon software engineering, Argon reached 77.9%. Google also said the model leads on benchmarks for finance, legal work, business automation, and long-video understanding, including 51.3% on Zapier’s Automation Bench and 91.7% on LV Bench.

Safety measures and internal debate

Google highlighted work on four safety areas before broader release: misuse prevention, prompt-injection resilience, monitoring for misalignment, and hardened training and testing environments. The company said it uses monitoring of model activations and reasoning traces, along with red-team testing and execution controls. At the same time, reports emerged that some employees had questioned Argon’s performance in internal testing, a claim Google disputed.

CONCLUSION

Gemini 4 Argon is Google’s strongest bid in months to regain ground at the high end of AI, combining frontier benchmark claims with an unusually cautious cyber-focused rollout. Its real standing will depend less on internal results than on whether outside developers and security teams find the model as capable in practice as Google says it is.

Ask a question

More from AI