Tech • AI • Robotics • Game

VIDEO
ENFR

Huge Gemini 4 Pro Leaks! GPT-6 Sol Testing, Grok 4.7 Update, Google RSI & More! AI News

9.5/10
AIWorldofAISeptember 18, 2026 at 06:15 AM20:59
Audio player
0:00 / 0:00

TL;DR

A wave of reported AI model tests and product updates suggests intensifying competition, led by a possible Google Gemini 4 Pro checkpoint, new signs of OpenAI GPT-6 Soul, and fresh releases or previews from xAI, Anthropic, Google DeepMind, and Xiaomi.

KEY POINTS

Gemini 4 Pro leak gains credibility

A new Google checkpoint reportedly being tested in public model arenas is increasingly believed to be Gemini 4 Pro. One clue is its apparent ability to sustain extremely large outputs, with a stress test generating highly detailed, non-repetitive data until a browser crashed, behavior that aligns with a reported 256K output token limit.

Strong coding and graphics outputs

Early examples attributed to the checkpoint show unusually polished code generation for visual tasks. Reported outputs include a Mario Kart-style game, a 3D Formula 1 car simulation, detailed animated SVG scenes, and console mockups with lighting, motion and layered structure that compare favorably with top current systems.

Animated console demos stand out

One standout prompt asked for an animated Nintendo Switch with controllers snapping into place and a boot sequence, while another produced a detailed PS5 controller and console in pure SVG. In side-by-side comparisons, the Gemini checkpoint was said to outperform some rivals on polish, though not every comparison placed it first.

GPT-6 Soul appears close

OpenAI is reportedly preparing GPT-6 Soul, a lower-tier sibling to GPT-6 Astra, with speculation around a near-term debut. The model is said to have an April 30 knowledge cutoff, matching Astra, and early tests suggest strong performance, richer formatting, improved memory behavior and better personalization in web app trials.

Soul shows strong value positioning

Reported evaluations suggest GPT-6 Soul can produce high-quality outputs at lower cost and latency than Astra while still approaching frontier-level capability in several domains. A cited 3JS generation of a 3D house environment was presented as a major step up over earlier Soul variants.

Claude Opus 5.2 surfaces in testing

Anthropic’s Claude Opus 5.2 was reportedly tested on multiple accounts, with early impressions pointing to strong results using fewer tokens than expected. That could signal a more efficient model, although wider benchmarks and public access have not yet arrived.

Grok 4.7 may be nearing release

xAI’s Grok 4.7 reportedly appeared in Google Cloud quotas and system limits, a sign that broader availability could be close. Early benchmark figures showed 6.4 on KernelBench Mega, ahead of Gemini 3.8 Flash at 2.7 and GPT-5.6 Soul at 2.6, but still behind several leading models including GLM 5.3, DeepSeek V4.1 Flash, GPT-6 Astra Pro and Claude Fable 5.

CUDA gains look modest but real

On KernelBench CUDA, Grok 4.7 posted 9.5, slightly above Grok 4.6 at 9.4. The early picture suggests a meaningful but not transformational upgrade, with improvements more incremental in some workloads than initial expectations may have implied.

Union Alpha turns out to be a routed system

A stealth model called Union Alpha, initially praised for near-frontier coding performance at much lower cost, was later identified as Pareto 269 from Unbiased. Rather than being a single model, it is a routing system that distributes work across multiple open and frontier models, helping explain why it appeared unusually close to state-of-the-art performance.

New decision model targets agents

A separate system called Jev, from Typesafe AI, takes a different approach from chatbots. It is designed for structured decision-making rather than text generation, returning calibrated outputs with probabilities, and is claimed to be 20 to 200 times faster and 40 to 400 times cheaper for certain routing, classification and action-selection workflows.

Google DeepMind pushes recursive self-improvement

Google DeepMind researchers introduced Dream RSI, a framework for recursive self-improvement in which an agent learns from earlier discovery attempts to improve future search strategies. The model weights do not change directly; instead, the exploration policy improves over time, with demonstrations spanning algorithms, GPU kernel engineering and model-development tasks.

Claude Code and Xiaomi add momentum

Claude Code has introduced Projects, enabling multiple tasks to run in parallel threads with shared memory and files, initially in beta for select Pro and Max users. Meanwhile, Xiaomi said Mimo 2.6 is in the middle of a large reinforcement-learning run, with roughly 2 billion tokens per training step across thousands of parallel rollouts and plans to share more details publicly.

CONCLUSION

The latest leaks, tests and product rollouts point to a faster-moving AI race in which model quality, efficiency, context size and agent capabilities are all advancing at once. If the reported results hold up, Google, OpenAI, Anthropic, xAI and Xiaomi are all preparing meaningful upgrades rather than incremental maintenance releases.

Ask a question
Full transcript

More from AI