Daily Podcast full article
Huge Google DeepMind RSI Leaks! GPT-6 Astra Nerfed, Kimi K2.8 Code, & More! AI News
A weekend burst of AI news has turned one theme into the center of the frontier race: systems that help build better systems. The latest reports connect Google DeepMind to alleged recursive self-improvement work, while Moonshot AI pushes Kimi K2.8 into coding workflows, OpenAI explains why GPT-6 Astra felt “nerfed,” and Anthropic argues that the whole industry may need to slow down before agentic systems outrun safety practice.

The story so far
The headline item is not a conventional model launch. It is the claim that Google DeepMind is moving deeper into recursive self-improvement, or RSI: the idea that AI systems can help evaluate, repair, train, or design later AI systems. The 8news subject page published on September 13 frames the leak trail around Google Vertex references to an “RSI model,” and connects that speculation to reports that Sergey Brin is pushing RSI-related resources while Demis Hassabis remains focused on AGI strategy .
That framing matters because RSI is not just another benchmark category. If an AI lab can use today’s models to accelerate the research, coding, evaluation, and debugging of tomorrow’s models, the competitive variable changes from “who has the best model today?” to “who can improve fastest?” In practical terms, that could mean agent loops that test model failures, propose training changes, improve tool-use scaffolding, write evaluation code, and feed results back into the next iteration.
But the current state still needs a careful verb: the public evidence points to leaks, model names, and community interpretation, not to a verified announcement that Google has achieved fully autonomous self-improvement. The September 13 8news report itself treats the Google RSI item as a leak-driven claim tied to Vertex and to broader DeepMind acceleration, rather than as a confirmed product release . The responsible reading is therefore: Google appears to be at least exploring or operationalizing RSI-like research loops, while the industry waits for externally testable proof.
Why Google’s RSI rumor hit a nerve
The reason the rumor spread quickly is that it fits the visible cadence of frontier AI. 8news notes that Google has recently shipped Gemini updates at an unusually fast pace, including closely spaced Gemini 3.7 and Gemini 3.8 releases, and that agentic loops for evaluation and refinement have become part of the public discussion around DeepMind’s roadmap .
The deeper question is whether this is simply better internal tooling or the beginning of a compounding research flywheel. A normal tool improves developer productivity. An RSI loop, even a partial one, improves the process that improves the model. That difference is why even vague evidence gets attention: once the loop closes tightly enough, progress can become less dependent on linear human research schedules.
Still, there are hard limits. AI can generate code, run tests, and inspect failures, but the bottlenecks in frontier model development also include data quality, evaluation validity, compute allocation, infrastructure reliability, safety review, and scientific judgment. If Google is building RSI infrastructure, the most plausible near-term form is not a rogue machine rewriting itself overnight. It is a large, managed research-and-engineering pipeline in which models increasingly act as junior researchers, test authors, code reviewers, and experiment runners.
Kimi K2.8: the quiet coding upgrade
While Google’s RSI rumor dominated attention, Moonshot AI delivered a more concrete update. IT之家 reported on September 11 at 16:00 Beijing time that Moonshot’s Kimi K2.8 Preview had fully rolled out to Kimi Code, keeping the same model ID, kimi-for-coding, so clients and third-party tools did not need configuration changes .
The important detail is not the odd version number. It is the positioning: K2.8 Preview is described as close to K3 in overall performance, with better thinking efficiency, upgraded coding and agent capabilities, and access to a 1M-token context window across all membership tiers . It also supports low, high, and max thinking-effort settings, with max as the default, and routes no-thinking requests from the K3 series and K2.8 Preview to a no-thinking K2.8 path .
That makes K2.8 look less like a prestige model and more like an operational model. K3 may remain the flagship reference point, but a cheaper or more efficient near-K3 coder can be more useful in daily engineering. Developers do not always need maximum reasoning; they often need a model that stops overthinking routine edits, fits the whole project in context, and responds predictably inside tools.
This is also where Kimi connects back to the RSI theme. Better coding agents are not only end-user products. They are also the machinery labs use internally to write tests, clean datasets, improve harnesses, and automate experiment management. A “coding model” is increasingly a research accelerator.
Astra’s “nerf” was an infrastructure story
OpenAI’s part of the weekend news was more defensive. Users had been comparing GPT-6 Astra’s launch-day outputs with later outputs and complaining that the model looked weaker, especially on complex visual or 3D-style generations. The accusation was simple: Astra had been quietly nerfed.
A September 12 BlockTempo report, citing OpenAI Codex product lead Thibault Sottiaux, says OpenAI identified three causes behind the perceived quality drop: older skills built for previous models were triggering too often, an opt-in context-management experiment caused early stops or replies to older messages, and some misconfigured engines degraded quality for long-tail traffic . The context experiment was estimated to have affected roughly 4,000 to 5,000 users, and OpenAI said the problematic pieces had been fixed or disabled .
This is important because it shifts the diagnosis from “the model got worse” to “the serving stack changed the experience.” In frontier AI, quality is no longer only a property of weights. It is also a property of routing, inference engines, context managers, skill files, tool policies, and product experiments layered around the model. A top-tier model can feel mediocre if the orchestration layer calls the wrong skill, truncates the wrong context, or routes a request through degraded infrastructure.
OpenAI also reset paid-user usage limits as part of the response, a move that functioned as both compensation and confidence management . The lesson for users is blunt: benchmark scores do not guarantee a stable product experience. The lesson for labs is even sharper: if customers perceive a downgrade, the burden of proof moves quickly to the provider.
Anthropic says the frontier needs pacing
The safety backdrop intensified at the same time. AP reported on September 12 that Anthropic CEO Dario Amodei said the AI industry should slow its fast-moving development so safety measures can catch up . He warned that without such a slowdown, AI systems could within six to 12 months become capable of leading a swarm of agents that could take over the internet .
Amodei’s proposal includes giving outside evaluators “ongoing, employee-like access” so they can monitor safety practices, and AP reported that Anthropic plans to do this itself, including desks, badges, and company laptops for evaluators . AP also reported that OpenAI’s Sam Altman quickly committed to a similar independent-evaluator model .
This lands directly on the RSI debate. If models are beginning to accelerate AI research itself, then safety evaluation cannot remain a release-day checklist. It has to move inside the development pipeline, where training runs, agent swarms, model behavior, and infrastructure changes can be inspected before they become public systems.
The bigger picture
Taken together, the week’s developments show the frontier moving in two directions at once. Google’s alleged RSI work points toward faster internal model improvement. Kimi K2.8 shows coding agents becoming more practical and efficient. OpenAI’s Astra incident shows how fragile model quality can be when product layers misfire. Anthropic’s pacing argument shows that safety leaders now see the acceleration mechanism itself as a governance problem.
The near-term conclusion is not that recursive self-improvement has fully arrived. It is that RSI-like workflows are becoming the organizing idea behind frontier competition. The labs are no longer only training models; they are building factories for model improvement. The winners may be the companies that can make those factories fast, reliable, auditable, and safe enough to trust.
Sources from the last 72 hours
- [1]Huge Google DeepMind RSI Leaks! GPT-6 Astra Nerfed, Kimi K2.8 Code, & More! AI NewsSep 13, 2026, 6:15 AM UTC
- [2]月之暗面 Kimi K2.8 Preview 模型全量上线 Kimi Code,综合性能接近 K3Sep 11, 2026, 8:00 AM UTC
- [3]OpenAI 修正 Astra 三個毛病,宣布今晚午夜前重置用量額度Sep 11, 2026, 7:20 PM UTC
- [4]Anthropic CEO Dario Amodei says AI industry needs to give safety measures time to catch upSep 12, 2026, 4:37 PM UTC
AI-generated article based on recent web research, then preserved as a dated editorial snapshot.

Comments
Be the first to comment.