
Tech • AI • Robotics • Game
Google has unveiled Gemini 4 Argon, a new flagship AI model focused on software engineering, enterprise work, and cyber defense, with early access limited to trusted security teams because of its ability to find and fix serious vulnerabilities autonomously.
Gemini 4 Argon is rolling out first through Fairwind, Google’s security initiative for a small group of trusted cyber defenders. The company said the model’s capabilities require a phased release, with participation in the US government’s voluntary pre-release review process before broader availability. Access is set to expand later to paid API customers, Google AI Ultra subscribers, enterprises, developers, and consumers, though no date was given.
Google set an introductory price of $2 per million input tokens and $10 per million output tokens. Cached input tokens, for reused context, are discounted by 95%, or about $0.10 per million. The use of “introductory” suggests pricing may rise after launch.
Google said Argon was trained to handle cyber defense tasks end to end, including finding critical vulnerabilities, validating them, and patching them autonomously. Because the same knowledge used to fix flaws can also be used to exploit them, a guardrail-free version is being provided only to trusted defenders and Google’s internal teams. Standard versions are designed to refuse assistance for cyberattacks and other high-risk misuse.
In an early demonstration with cloud security company Wiz, Argon identified a critical flaw in healthcare software used by hospitals worldwide that exposed sensitive personal data. Google said earlier frontier models had missed the issue. The result is being presented as evidence that the new model can detect vulnerabilities beyond prior systems.
On CWE Bench V1, which measures how well models remediate software vulnerabilities, Argon scored 68%, tying for first place against systems including GPT-6 Astra, Claude Opus 5.5, and Grok 4.7. Google also said Argon outperformed its earlier security model, Gemini 3.8 Flash Cyber, on internal tests and on Wiz black-box penetration benchmarks involving live web targets without source-code access.
Google said thousands of employees have already been using Argon for coding, research, and writing. The company described applications ranging from debugging and large code migrations to algorithm design. In quantum computing work, Argon reportedly improved a published baseline by 40% in minutes by optimizing subroutines that reduce “space-time resources,” a measure combining qubits and gate operations.
Argon agents are being used to migrate code from C and C++ to Rust, a language designed to prevent classes of memory-safety bugs. Google said projects range from libraries such as RE2 and libgav1 to the Fuchsia Zircon kernel, which has more than 800,000 lines of code. In another internal effort, Argon-driven memory optimizations across Google’s data center fleet freed more than 300 tebibytes of memory, with total savings estimated between 500 tebibytes and 1 pebibyte.
In work on libgav1, Google’s open-source video decoder, Argon agents replaced 32,000 lines of hand-tuned SIMD code with safe Rust structured for compiler vectorization. Google said the result was a memory-safe decoder that ran 2.7 times faster than the prior Rust port while preserving identical video output and narrowing the gap with the optimized C++ version.
Google raised Argon’s output limit to 1 million tokens, up from 64,000, to support long, multi-step workflows. On DeepSuite V1.1, a benchmark for long-horizon software engineering, Argon reached 77.9%. Google also said the model leads on benchmarks for finance, legal work, business automation, and long-video understanding, including 51.3% on Zapier’s Automation Bench and 91.7% on LV Bench.
Google highlighted work on four safety areas before broader release: misuse prevention, prompt-injection resilience, monitoring for misalignment, and hardened training and testing environments. The company said it uses monitoring of model activations and reasoning traces, along with red-team testing and execution controls. At the same time, reports emerged that some employees had questioned Argon’s performance in internal testing, a claim Google disputed.
Gemini 4 Argon is Google’s strongest bid in months to regain ground at the high end of AI, combining frontier benchmark claims with an unusually cautious cyber-focused rollout. Its real standing will depend less on internal results than on whether outside developers and security teams find the model as capable in practice as Google says it is.
Ask a question