
Tech • AI • Robotics • Game
Google’s Gemini 4 Argon has entered limited release with unusually aggressive pricing, a 1 million-token output cap, and benchmark results that position it as a strong contender for everyday multimodal and long-horizon AI work.
Gemini 4 Argon is currently available only to selected cybersecurity partners and trusted testers through Google DeepMind’s Fairwind program. A broader launch is expected for paid API customers and Google AI Ultra subscribers, signaling a phased commercial release rather than an immediate public debut.
Introductory pricing is set at $2 per 1 million input tokens and $10 per 1 million output tokens, with cached input priced 95% lower. After the promotion, pricing is expected to move to $4 input and $20 output per million tokens, still placing the model among the more aggressive frontier offerings on cost.
The most eye-catching technical change is the increase in maximum output from 64,000 to 1 million tokens. That scale could make Argon notably useful for long-running agentic tasks, software workflows, and extended reasoning chains that require sustained context and large final outputs.
On SWE-bench 1.1, Argon reportedly scored 77.9%, described as a state-of-the-art result for real-world long-horizon software engineering. It also led on BAI’s index spanning finance, coding, legal, and tax work, and ranked first on Automation Bench with 51.3% for end-to-end business workflows.
Argon also posted a reported 91.7% on LVB, setting a high-water mark for long-video understanding. That result reinforces Google’s broader positioning of Gemini as a multimodal system rather than a model focused only on text or code.
On the Artificial Analysis Intelligence Index, Argon scored 53, matching GPT-6 Astra and sitting just one point behind GPT-6.1 Soul. Yet its estimated cost was cited at about $1.99 per intelligence-index task, compared with $3.26 for Astra and $5.98 for Opus 5.5, suggesting unusually strong value even if it is not the outright top model on every benchmark.
One of the most notable claims is a 15% hallucination rate, described as the lowest measured by Artificial Analysis for a model scoring above 45 on its intelligence index. For comparison, Astra was cited at 51% and GPT-6.1 Soul at 54%, a gap that could matter for enterprise use where reliability is often more important than peak benchmark performance.
Argon reportedly reached 77.5% on Automation Bench, ahead of Sonnet 5.5 at 71.3%, and showed substantial gains on terminal-style evaluations. Even so, it is not being positioned as the best model for every coding or computer-use task, with some rivals still seen as stronger in specialized developer workflows.
Testers have highlighted detailed outputs including voxel-built environments, SVG rendering, simulated 3D scenes, dynamic lighting changes, and polished interactive demos. Examples such as a PS5 controller SVG, a mechanical bee, a coliseum scene, and physics-based 3D world building point to improvements in front-end generation and multimodal design quality.
Reporting has suggested Argon is larger than previous Gemini 4 models, indicating a possible increase in model scale alongside gains in reasoning and autonomy. Google also says the system is already assisting internal researchers and engineers with algorithmic and technical work, underscoring how AI models are increasingly being used to help build their successors.
Gemini 4 Argon appears to strengthen Google’s position through a mix of long-context capability, low hallucination rates, and unusually competitive pricing. If those metrics hold in wider release, it could become one of the most practical frontier models for daily multimodal and enterprise use.
Ask a question