Daily Podcast full article
Gemini 4 RSI Leak Destroys Astra and Fable in Benchmarks
A mysterious Arena entry labeled “gemini-3.8-flash” has reignited speculation that Google is testing a Gemini 4 Pro-class model under a familiar name. The leaked benchmark card says it beats GPT-6 Astra and Claude Fable 5.1, while newer official Android Bench data offers a useful reality check: the public Gemini 3.8 Flash is not the same thing as the rumored heavyweight.

The leak is really a naming mystery
The current story is not simply “Gemini 3.8 Flash is suddenly the best model in the world.” It is stranger than that. Reports circulating on September 18 describe an Arena model shown as gemini-3.8-flash but behaving, according to testers, unlike the speed-optimized Flash tier that Google had already put into the market earlier in the month . That mismatch is the entire reason the rumor escalated from another leaderboard screenshot into a possible Gemini 4 Pro leak.
The strongest version of the claim comes from Chinese AI coverage republished by Pedaily from Xinzhiyuan: a new Arena entrant carrying the gemini-3.8-flash label is “suspected” to be Gemini 4 Pro, and a widely shared scorecard puts it ahead of GPT-6 Astra and Claude Fable 5.1 in coding, agentic, reasoning and computer-use tasks . A more cautious Qiniu write-up, also published on September 18, separates what is confirmed from what is not: the confirmed part is the online buzz around an Arena entry and developer demos; the unconfirmed part is the model identity, the leaked benchmark table, pricing, memory and context-window claims .
That distinction matters. If the Arena system is just the public Gemini 3.8 Flash, then the leak is hype. If it is a masked checkpoint, then Google may be doing blind comparative testing under a safe label, a common enough tactic to reduce brand-driven voting bias. The public evidence so far supports only the weaker statement: testers saw an unusually capable gemini-3.8-flash entry and inferred that it might be Gemini 4 Pro .
What the leaked benchmark card says
The benchmark card driving the headline lists four claimed wins for “Gemini 4 Pro” over GPT-6 Astra and Claude Fable 5.1. Atoms’ September 18 summary reproduces the leaked table as follows: DeepSWE v1.1 at 88.7% for Gemini 4 Pro versus 86.9% for Astra and 69.1% for Fable 5.1; GDPval-AA v2 at 2064 Elo versus 1994 and 1853; Terminal-bench 2.1 at 95.3% versus 94.1% and 92.8%; and OSWorld-2.0 at 86.8% versus 84.5% and 77.9% .
Those numbers explain why the story spread so quickly. The alleged win over Astra is narrow in some places but symbolically important: software engineering, terminal use and OS-level computer operation are precisely the categories where frontier labs have been trying to prove real economic usefulness. The alleged gap over Fable 5.1 is much larger on DeepSWE and OSWorld, which makes the leak sound less like an incremental Google catch-up and more like a step-change .
But the same source stresses that neither the results nor the listed prices have been independently verified . That caveat should sit next to every reading of the leak. A screenshot without a reproducible harness, model hash, sampling settings, tool environment and contamination controls is not the same as a public benchmark result. It is a signal; it is not yet a settled ranking.
Why the demos were more persuasive than the numbers
The benchmark card was not the only reason developers paid attention. The September 18 reports describe a cluster of visual and interactive demos: SVG drawings, a graphite-style website that draws itself as the user scrolls, a pelican riding a bicycle with controls, a voxel pagoda, a helicopter model and small playable game-like outputs . These demos matter because they show behavior that users can evaluate directly, even if they cannot audit the benchmark pipeline.
They also fit the core anomaly. A fast Flash model should usually answer quickly, trading some depth for latency. The rumored Arena entry reportedly spent far longer on certain generations and produced unusually polished code-heavy visual artifacts . In other words, the argument is behavioral: the model’s pacing and output quality did not feel like a normal Flash-tier endpoint.
Still, “better than expected” does not prove “Gemini 4.” It could be an experimental 3.8 variant, a high-reasoning configuration, a harness difference, a temporary route to a stronger backend, or a future model hidden behind an old name. The most responsible interpretation is that the Arena label is not sufficient identification.
The RSI angle: powerful narrative, weak evidence
The “RSI” part of the headline is also doing a lot of work. Pedaily’s Xinzhiyuan-derived report says rumors point to Google having implemented recursive self-improvement internally, allowing Gemini 4 Pro to compete directly with Astra and Fable . Qiniu’s more careful account says the RSI loop claim has no official or paper-level proof and should be treated as rumor, even though Google-related discussions of recursive self-improvement and AI-assisted research have increased the narrative pressure around Gemini 4 .
That is the right framing. RSI, or recursive self-improvement, is the idea that an AI system helps improve the methods, tools or models that produce future AI systems. If a lab really operationalized it at scale, the market would care. But the leap from “Google is interested in RSI” to “this Arena model is an RSI-produced Gemini 4 Pro” is too large for the evidence available on September 19.
The story is therefore best read as two overlapping claims. Claim one: an Arena model labeled gemini-3.8-flash appears unusually strong and may be a masked Gemini 4-class checkpoint. Claim two: its strength comes from internal RSI. The first claim has community observations and repeated coverage behind it; the second remains speculative.
The official benchmark reality check
A major complication arrived from Google’s own Android Bench 2.0 update. The Android Developers Blog announced on September 17 that Android Bench 2.0 now includes long-horizon tasks, agent evaluations and continuous scoring, with newly added models including Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3 and Qwen 3.8 Max . In that official context, OpenAI’s GPT-6 Astra topped the new benchmark with a 28% pass rate .
9to5Google’s coverage of the same update underlined the contrast: Google had rated Gemini 3.7 and 3.8 Flash alongside GPT-6 and Fable 5.1, and GPT-6 Astra sat at the top with a 28% pass rate, while the benchmark focused on harder long-horizon Android development tasks rather than simpler incremental changes . That does not disprove the leak, because the leaked model may not be the public Gemini 3.8 Flash tested in Android Bench. But it does prevent an easy conclusion that “Gemini 3.8 Flash destroys Astra and Fable” across public, reproducible settings.
This is the key editorial point: the leaked story and the official benchmark can both be true if they refer to different systems. Public Gemini 3.8 Flash may trail Astra in Android Bench 2.0, while an unreleased Gemini 4 Pro checkpoint may be circulating under a confusing Arena alias. The confusion is the story.
Why Google would use a familiar alias
If Google is testing a stronger system under the gemini-3.8-flash name, the logic is not hard to imagine. A familiar alias gives the model live comparisons without announcing a product, avoids immediate expectation-setting around Gemini 4, and may reduce the halo effect that would occur if users saw “Gemini 4 Pro” in the interface. Arena-style testing is most useful when users judge outputs, not brand names.
There is also a roadmap reason. The September 18 coverage repeatedly frames the leak against months of developer impatience for a new Pro-class Gemini release . If Google has a model that can seriously challenge Astra and Fable, blind testing would let it gather preference data and stress-test tool behavior before launch. If it does not, the aliasing also gives Google room to experiment without turning every checkpoint into a public promise.
Bottom line
As of September 19, the safest headline is not that Google has officially launched Gemini 4 Pro. It has not. The safer headline is that a model labeled gemini-3.8-flash has generated credible curiosity because its reported behavior, demos and leaked scorecard look too strong for a normal Flash-tier release .
The “destroys Astra and Fable” claim belongs to the leaked benchmark card, not to an independently verified public leaderboard. The official Android Bench 2.0 data actually puts GPT-6 Astra ahead in a difficult Android development setting . But if the Arena entry is indeed a masked Gemini 4 Pro checkpoint, then the leak signals something important: Google may be closer than expected to re-entering the frontier-model fight not with a faster Flash, but with a hidden heavyweight wearing a Flash badge.
Sources from the last 72 hours
- [1]刚刚,Gemini 4 Pro偷跑上线,碾压Astra和Fable_投资界Sep 18, 2026, 3:46 AM UTC
- [2]Gemini 4 Pro偷跑上线?匿名现身Arena,网传基准碾压Astra和Fable(2026年9月) | 七牛云Sep 18, 2026, 12:00 AM UTC
- [3]Gemini 4 Pro Leaks: The Lasted SOTA Model?Sep 18, 2026, 12:00 AM UTC
- [4]Android Bench 2.0: pokonywanie kolejnych granic dzięki wymagającym zadaniom długoterminowymSep 17, 2026, 12:00 AM UTC
- [5]Android Bench 2.0 focuses on long-horizon tasks, agent evaluationsSep 17, 2026, 4:00 PM UTC
AI-generated article based on recent web research, then preserved as a dated editorial snapshot.

Comments
Be the first to comment.