Tech • AI • Robotics • Game

VIDEO
ENFR

Gemini 4 Argon Is Google’s Most Powerful AI Model + Early Tests!

9.4/10
AIWorldofAISeptember 30, 2026 at 09:53 PM14:13
Audio player
0:00 / 0:00

TL;DR

Google’s Gemini 4 Argon has entered limited release with unusually aggressive pricing, a 1 million-token output cap, and benchmark results that position it as a strong contender for everyday multimodal and long-horizon AI work.

KEY POINTS

Limited rollout begins

Gemini 4 Argon is currently available only to selected cybersecurity partners and trusted testers through Google DeepMind’s Fairwind program. A broader launch is expected for paid API customers and Google AI Ultra subscribers, signaling a phased commercial release rather than an immediate public debut.

Pricing undercuts many rivals

Introductory pricing is set at $2 per 1 million input tokens and $10 per 1 million output tokens, with cached input priced 95% lower. After the promotion, pricing is expected to move to $4 input and $20 output per million tokens, still placing the model among the more aggressive frontier offerings on cost.

A major jump in output length

The most eye-catching technical change is the increase in maximum output from 64,000 to 1 million tokens. That scale could make Argon notably useful for long-running agentic tasks, software workflows, and extended reasoning chains that require sustained context and large final outputs.

Strong benchmark showing across business and software tasks

On SWE-bench 1.1, Argon reportedly scored 77.9%, described as a state-of-the-art result for real-world long-horizon software engineering. It also led on BAI’s index spanning finance, coding, legal, and tax work, and ranked first on Automation Bench with 51.3% for end-to-end business workflows.

Multimodal and video performance stand out

Argon also posted a reported 91.7% on LVB, setting a high-water mark for long-video understanding. That result reinforces Google’s broader positioning of Gemini as a multimodal system rather than a model focused only on text or code.

Cost-to-performance may be its main selling point

On the Artificial Analysis Intelligence Index, Argon scored 53, matching GPT-6 Astra and sitting just one point behind GPT-6.1 Soul. Yet its estimated cost was cited at about $1.99 per intelligence-index task, compared with $3.26 for Astra and $5.98 for Opus 5.5, suggesting unusually strong value even if it is not the outright top model on every benchmark.

Low hallucination rate is a differentiator

One of the most notable claims is a 15% hallucination rate, described as the lowest measured by Artificial Analysis for a model scoring above 45 on its intelligence index. For comparison, Astra was cited at 51% and GPT-6.1 Soul at 54%, a gap that could matter for enterprise use where reliability is often more important than peak benchmark performance.

Agentic gains, but not a universal leader

Argon reportedly reached 77.5% on Automation Bench, ahead of Sonnet 5.5 at 71.3%, and showed substantial gains on terminal-style evaluations. Even so, it is not being positioned as the best model for every coding or computer-use task, with some rivals still seen as stronger in specialized developer workflows.

Early demos suggest stronger visual and interactive generation

Testers have highlighted detailed outputs including voxel-built environments, SVG rendering, simulated 3D scenes, dynamic lighting changes, and polished interactive demos. Examples such as a PS5 controller SVG, a mechanical bee, a coliseum scene, and physics-based 3D world building point to improvements in front-end generation and multimodal design quality.

Google appears to be scaling both model size and internal use

Reporting has suggested Argon is larger than previous Gemini 4 models, indicating a possible increase in model scale alongside gains in reasoning and autonomy. Google also says the system is already assisting internal researchers and engineers with algorithmic and technical work, underscoring how AI models are increasingly being used to help build their successors.

CONCLUSION

Gemini 4 Argon appears to strengthen Google’s position through a mix of long-context capability, low hallucination rates, and unusually competitive pricing. If those metrics hold in wider release, it could become one of the most practical frontier models for daily multimodal and enterprise use.

Ask a question

More from AI