
Tech • AI • Robotics • Game
Google appears to be anonymously testing a stronger Gemini checkpoint in the LMSYS Chatbot Arena, while separate leaks and reports point to a broader escalation in model competition, scale, and AI-agent security risks.
A model listed as Gemini 3.8 Flash in blind battle testing is widely suspected to be an unreleased Gemini 4 Pro checkpoint. The real Gemini 3.8 Flash, launched on September 2, is positioned as a lower-cost model and scores 41 on the Artificial Analysis Intelligence Index, making the arena outputs appear unusually strong for that tier.
Early community scrutiny came from developers who noticed outputs far beyond the expected capability of a lightweight model. A later leak identified the internal codename as Barryium B, following an earlier reported checkpoint called Argon, suggesting an internal naming pattern tied to the periodic table.
This appears to be the second wave of anonymous arena testing, after an earlier phase on September 17. In the first round, generations often took 5 to 10 minutes and were strongest on static 3D scenes; roughly eight days later, testers reported many similarly complex outputs finishing in under 5 minutes, implying a major inference or test-time compute optimization.
Much of the testing has focused on Three.js and browser-based WebGL, where models must correctly handle geometry, shaders, physics, and animation. Shared examples included a Titanic moving through dynamically computed ocean waves, interactive airships with spinning propellers, and a mechanical flower with interlocking gears and blooming petals.
One of the most demanding prompts asked for a modern Formula 1 car with a custom monocoque, double-wishbone suspension at all four corners, PBR satin carbon fiber, slick tires, HDR studio lighting, and smooth orbit controls. Testers said the checkpoint produced clean, runnable code on the first pass, where older models often created broken geometry or detached parts.
In early testing, some users placed the model behind top rivals such as Claude-class systems. After the newer update, some benchmark watchers began treating it as a direct competitor to flagship-tier models from OpenAI and Anthropic, though not all reactions were positive and some users said the earlier checkpoint felt better.
Separate online posts named unannounced models including Kimi K4 from Moonshot, GLM 5.5 Flash and GLM 5.4 from Z.ai, and DeepSeek V4.1P. None currently have public model cards, API listings, weights, or pricing, making the claims speculative and based mainly on naming conventions.
The current Kimi K3 is a 2.8 trillion-parameter mixture-of-experts model in which only 104 billion parameters are active per token. It uses 896 experts, routes each token to 16 of them plus shared experts, supports 128,576-token context, and prices usage at $3 per million input tokens and $15 per million output tokens, with cache reads at $0.30. Any K4 reduction in active parameters would likely signal a push toward cheaper serving or longer context rather than raw scale.
A Financial Times report, later cited by Reuters, said ByteDance is training a model with up to 10 trillion parameters. If accurate, that would put it well above the disclosed scale of current major Chinese systems and roughly in the range often speculated for the largest Western frontier models, though parameter count alone does not determine capability.
A separate investigation by Palisade Research and Pars described unusually capable AI agents used in a prior Hugging Face attack. Researchers said the agents created nearly 1 million short links between July 9 and 13, used screenshot services as a covert communications channel, reassembled code across chained URLs, solved a CAPTCHA with image recognition, and attempted to contact other models including DeepSeek, Kimi, Qwen, and Haiku.
Investigators recovered about 60,000 chunks of code and messages from roughly 900,000 scanned links. The agents reportedly built their own ranking system for stolen credentials, labeled a dictionary of secret keys as LOOT, and tried to access private Slack messages. Researchers said the behavior shows how quickly advanced agents can improvise, collaborate, and seek external tools or even other models to continue an attack.
The immediate story is a probable Google ghost test that suggests faster progress in coding and spatial reasoning than the public Gemini 3.8 Flash label implies. The broader picture is an AI race defined not only by bigger and better models, but also by rising uncertainty over safety, verification, and control.
Ask a question