
Tech • AI • Robotics • Game
Google is internally testing a new Gemini 4 coding model codenamed Carbon that employees say approaches Anthropic’s Claude Opus 5.5, while Microsoft has launched a separate low-cost AI system built specifically for fast, structured decision-making.
Internal materials reviewed by Business Insider indicate that a new Gemini 4 checkpoint called Carbon was recently enabled in Jet, Google’s internal coding platform linked to its broader autonomous coding efforts. Several employees described the model as highly capable, with one saying it felt comparable to Opus 5.5 for coding. Others in internal chat reportedly called it “really good” and said the latest Gemini Pro next model felt strong in testing.
The praise reflects a shift from only weeks earlier, when early versions of Argon, Google’s newly announced flagship model, were reportedly seen as trailing on complex coding tasks. One employee had compared that earlier system to Claude Opus 5, suggesting Google lagged a generation behind in the coding performance most valued by developers. If Carbon is genuinely near Opus 5.5, it would mark a meaningful narrowing of that gap.
Internal documents show Google has tested a series of Gemini 4 variants named Argon, Barium, and Carbon. One document said the internal model Barium B was selected for public release under the Argon label, meaning the public branding does not map cleanly to internal development names. Carbon is described as a checkpoint, or mid-training snapshot, and it remains unclear whether it will ship as an Argon update or as a separate model.
Google announced Argon on September 30, highlighting frontier-level results in coding, knowledge work, legal and finance tasks, and defensive cybersecurity. Access initially went to partners in the Fairwind testing program, with broader rollout to paid API customers and Google AI Ultra subscribers expected later. Regular users still have not received it, leaving rivals such as OpenAI and Anthropic with publicly accessible models already in the market.
Google has shown charts claiming Argon outperforms top competitors, but outside users have not been able to verify those claims in real-world use. That matters because Gemini models have faced criticism for looking dominant on benchmark slides while underwhelming after release. The gap between internal enthusiasm and public availability has sharpened scrutiny over whether Google can translate lab performance into products.
External testing watchers have spotted Gemini 4 Argon appearing in Anti-Gravity, Google’s coding tool, with context windows of 256,000, 512,000, and 900,000 tokens. The higher settings reportedly consume about 1.3 times and 1.8 times more quota per turn. Google’s web app has also added low, medium, and high “thinking effort” settings, while a workplace-oriented universal Gemini agent and a previously surfaced feature called Chief of Staff point to a heavier focus on autonomous software agents.
Some internal discussion has reportedly linked Carbon’s rapid progress to recursive self-improvement, the idea that an AI system helps refine its own coding and problem-solving abilities across repeated cycles. However, that claim does not appear in the core Business Insider reporting itself and remains unverified. Even if the technique is being used, matching a rival model that has already been public for weeks would not, by itself, prove runaway self-improvement.
While Google pushes toward larger, more autonomous coding models, Microsoft on October 9 released Decision One, a small model optimized for choosing among predefined answers such as yes-or-no, multiple choice, rankings, and rubric-based grading. Instead of generating long responses, it returns structured probabilities that software can act on directly for routing, classification, prioritization, verification, and workflow control.
Microsoft said Decision One was post-trained from Alibaba’s Qwen 3.5 9B and topped 36 benchmarks spanning nearly 150,000 questions. The company said it was 2.5 times faster than the nearest runner-up and roughly 35 times faster than GPT-6 Soul on median response time. It also flipped answers only 1.3% of the time when prompts were rewritten, while input pricing was set at $0.042 per million tokens and output tokens were free.
Internal teams across Microsoft are using the model for tasks such as sorting Xbox user feedback, evaluating Copilot responses, assisting incident response, and guiding scientific replanning loops in Microsoft Discovery. In some cases, the company said it matched or beat larger language models while running at 14 times the speed, 100 times faster, or 200 times lower cost depending on the task. The model is available through Microsoft Foundry and OpenRouter.
The latest moves show two distinct AI strategies taking shape: Google racing to regain credibility in frontier coding, and Microsoft targeting the economics and control layer that practical AI agents need. Whether Carbon reaches users in its current form may determine how quickly Google can turn internal momentum into a public comeback.
Ask a question