
Tech • AI • Robotics • Game
A model listed on Arena as Gemini 3.8 Flash is prompting speculation that Google is quietly testing a far more capable unreleased system, possibly tied to Gemini 4.
A model appearing on the blind test platform Arena carried the label Gemini 3.8 Flash, a name already used for a public Google model released on September 2. What drew attention was not the alias itself, since anonymous testing is common, but that this system behaved nothing like a speed-focused Flash tier. On a prompt to produce a raw SVG render of a PlayStation 5, it reportedly paused for about 10 minutes before returning highly detailed vector code, far beyond the quick, approximate output expected from the public version.
Developers circulated examples showing unusually precise SVG, voxel and 3D outputs, including renders of a PS5, BMW M4, and other complex objects. Side-by-side comparisons with older Gemini checkpoints suggested a categorical improvement in geometry, shading, alignment and depth rather than a small iterative gain. Tests such as a pelican riding a bicycle under a starry twilight sky reportedly showed accurate placement, clean shapes and coherent multi-part execution in a single pass.
The long wait times are fueling a theory that the system is doing extensive internal planning before producing any tokens. Instead of streaming rough answers quickly, the model appears to spend minutes writing, refining and debugging before replying. That pattern fits a high-compute reasoning system more than a lightweight Flash product, suggesting the label may be camouflage for a larger checkpoint.
Showcase builds attributed to the model include a pencil-textured website where scrolling darkens lines like a brush stroke, a cyberpunk-retrofuturist interface with 3D elements, and an Airbus H145 helicopter model produced in about 10 minutes. More advanced demos include a Minecraft-style build, a 3D kart racer and a flight simulation scene. The recurring claim is that the model handles broad creative direction and technical execution together, with less iterative correction than rival systems typically need.
Unverified benchmark charts circulating online place the mystery model at roughly 88% on SWEV1.1, 264 ELO on GPQA Val A v2, 95.3% on Terminal Bench 2.1, and 86.8% on OSWorld 2.0. If accurate, those scores would put it ahead of named competitors such as Astra and Fable 5.1 across coding agents, reasoning and computer-use tasks. None of those figures has been confirmed by Google.
Another claim drawing attention is possible pricing of $2.25 per million input tokens and $11.25 per million output tokens. That would be far below the $10 input and $50 output rates associated with some competing frontier models, while still costing about three times more than the promotional rates for public Gemini 3.8 Flash. That gap reinforces the view that the Arena model may belong to a higher-tier product.
Researchers tracing routed calls have claimed the model exposes a 10 million-token input ceiling and 256,000-token output ceiling. Those numbers are far beyond the public Flash configuration of 1,048,576 input and 65,536 output tokens. Reports also point to persistent cross-session memory, direct internet access, sandboxed execution for untrusted code and even native robotic motor-control support, though all remain unverified.
The timing aligns with a difficult year for Google’s top-end model lineup. Gemini 3.1 Pro arrived on February 19, while a previewed Gemini 3.5 Pro reportedly slipped and was later said to have been canceled. Google has publicly said it began its most ambitious pre-training effort yet and that Gemini 4 is important to staying competitive at the frontier, making the Arena model a plausible release candidate.
Some of the speculation centers on RSI, or recursive self-improvement, as a possible factor in the model’s progress. Recent comments from Google DeepMind leadership and new research such as Dream RSI show that self-improving agent loops are central to current strategy. There is no proof that such methods accelerated Gemini 4, but the rumors have amplified interest because of what they could imply for model capability growth.
The strongest public evidence so far suggests that the model labeled Gemini 3.8 Flash on Arena is not the public Flash system. If Google is using that slot to test a hidden frontier model, the reveal could signal a major shift in the AI pricing and capability race.
Ask a question