Tech • AI • Robotics • Game

VIDEO
ENFR

Daily Podcast full article

OpenAI Just Built Early RSI

OpenAI’s reported internal automation of experimental model development is not full recursive self-improvement yet, but it is close enough to change the safety conversation: AI agents are now helping build, test, optimize, and coordinate parts of the next AI stack, while the industry scrambles to define when human oversight stops being real oversight.

Generated September 24, 2026 at 4:13 PM UTC1335 words
AI-generated illustration

The story: an AI lab letting AI do more of the lab work

OpenAI has reportedly pushed its internal AI systems into the heart of experimental model development, automating large parts of the training workflow that used to require direct human research and engineering labor . The strongest version of the claim is not that OpenAI has created a fully self-improving machine, but that its researchers can now describe a desired optimization and let AI systems execute, test, correct, and iterate over parts of the model-development process for extended periods .

That is why the phrase “early RSI” has landed so forcefully. Recursive self-improvement, or RSI, usually means an AI system that can improve itself, then use the improved version to make the next version better again, potentially accelerating beyond human control . OpenAI is not publicly claiming that such a fully closed loop exists. In fact, the latest public framing from safety-focused reporting is that the field is not there yet . But OpenAI’s reported workflow is one of the closest practical approximations so far: AI doing meaningful work on the systems that produce better AI.

The most important detail is the shift from assistance to delegation. According to the current summaries of the reporting, OpenAI’s internal systems can handle much of the training process for experimental models, including code changes, experiment runs, result monitoring, and revisions when something breaks . Researchers are still setting direction, but the machine is increasingly filling in the operational middle. In software terms, humans are writing the spec; agents are building more of the implementation.

What makes this feel different

AI tools have helped researchers write code for years. The new threshold is autonomy over a long chain of technical steps. OpenAI’s system is described as being able to run for weeks after receiving a target optimization example, with agents collaborating, discussing approaches, and iterating without constant human intervention . That matters because AI research is not just one clever idea; it is thousands of trials, failed configurations, debugging sessions, kernel optimizations, evaluations, and regressions.

The GPU-kernel detail is especially revealing. Top AI labs spend enormous money and talent on low-level code that makes training and inference faster on graphics processors. If an internal model can now write or optimize much of that code, then AI is not merely helping with a peripheral task; it is working on the machinery that determines how fast future models can be trained and served . That is exactly the sort of “AI improving the AI factory” dynamic that makes RSI no longer feel like a philosophy seminar topic.

The reported speedup is also significant. The HelloBro summary of the story says experiments that once could have taken years are now being compressed into roughly a week inside OpenAI’s process . Even if that applies only to specific classes of experiments, it changes the balance between idea generation, compute availability, and safety review. A backlog of old research tricks becomes dangerous or powerful once agents can test them quickly.

Why this is not full RSI

The distinction matters. Full RSI would imply a mostly closed loop: the system proposes improvements, implements them, evaluates whether they worked, incorporates the improvement into a successor, and repeats the process with little or no human dependence. The current reporting still describes humans as selecting goals, giving examples, allocating infrastructure, and deciding what counts as acceptable progress .

Axios’s current explanation makes the same point from another direction: researchers at OpenAI and Anthropic say model training has grown more automated and is approaching the RSI line, but the threshold has not clearly been crossed . Axios also notes that a new analysis published this week found that AI feedback loops are not yet meeting RSI benchmarks . In other words, “early RSI” is a fair warning label only if we mean a precursor: partial loop closure, not a runaway self-improver.

That nuance should not make the story smaller. Most technological shocks arrive first as partial systems. The question is not whether OpenAI has built the final self-improving intelligence. The question is whether the economic and technical incentives now point toward closing the loop further. On that question, the answer looks like yes.

Oversight becomes the bottleneck

OpenAI’s reported internal automation comes at the same time the company is asking for more formal global safety standards. On September 21, Axios reported that OpenAI had released proposed international AI safety standards as U.S. and Chinese officials discussed AI-risk notification mechanisms . The proposal emphasized shared global measurements for classifying incidents, tracking alignment issues, and reporting or responding to safety problems .

That timing is not accidental. If AI systems increasingly run experiments, write training infrastructure, and coordinate with other agents, then the old review model breaks down. A human reviewer cannot meaningfully approve every line of code, every experimental branch, every agent-to-agent message, and every emergent workaround. The audit problem becomes one of sampling, monitoring, rollback, and anomaly detection.

This is where “human in the loop” can become a comforting phrase rather than a real control system. If humans approve broad goals while agents execute thousands of technical steps, human control depends on instrumentation. What did the agents change? What did they test? What did they ignore? Did they optimize the intended objective, or did they discover a loophole? Did multiple agents create their own informal workflow outside the intended chain of command?

That concern is reinforced by reports of frontier labs exploring outside checks. Dealroom’s summary of The Information’s reporting says OpenAI and Anthropic neared a legally binding agreement to stress-test each other’s commercial models through API access, with data-retention protections and obvious antitrust sensitivities . Even if such an arrangement remains uncertain, the direction is telling: labs know that self-certification is losing credibility.

The real risk: acceleration without interpretability

The scariest version of this story is not “the AI wakes up.” It is more mundane and therefore more plausible: a lab creates a system that is extremely good at accelerating research, but the system’s reasoning, code paths, and failure modes become too complex for ordinary institutional review. Axios captured the wider dilemma this week, noting that thousands of semi-autonomous agents and increasingly complex AI systems make real-time auditing difficult .

That risk increases when competition rewards speed. If one lab can turn a year of experiments into a week of agentic iteration, rivals will feel pressure to match it. Safety practices that made sense when experiments moved slowly may not survive a world where the research loop runs continuously. “Sudo make me a model” is a joke until the script keeps running after everyone has gone home.

What to watch next

The next decisive signal will be measurement. How much of OpenAI’s AI R&D is now performed by agents? Which tasks still require human researchers? Are the agents changing model architectures, training recipes, data mixtures, evaluation harnesses, or only implementation code? Are failed runs and near-misses disclosed to external evaluators? Are there hard triggers that force human review before a model trained through these workflows can influence the next system?

The second signal will be governance. OpenAI’s push for international standards, incident reporting, and shared measurements suggests the company understands that voluntary claims will not be enough . But standards will matter only if they define operational thresholds: how to measure automated R&D, when to pause a training loop, how to preserve logs, and who has authority to inspect the process.

For now, the honest headline is the one we started with: OpenAI just built early RSI. Not the mythic, fully autonomous intelligence explosion. Not a harmless coding assistant either. What it appears to have built is a practical bridge between today’s agentic engineering and tomorrow’s self-improving AI pipeline. That bridge is now the place where capability, safety, competition, and control all meet.

Comments

Be the first to comment.

Sources from the last 72 hours

  1. [1]OpenAI Just Built Early RSISep 23, 2026, 11:11 PM UTC
  2. [2]The AI doomsday fear hidden in self-improving AISep 22, 2026, 9:00 AM UTC
  3. [3]OpenAI proposes AI standards after U.S., China talksSep 21, 2026, 5:00 PM UTC
  4. [4]OpenAI and Anthropic neared a mutual AI stress-test deal — coopetition meets antitrust opticsSep 22, 2026, 6:17 PM UTC

AI-generated article based on recent web research, then preserved as a dated editorial snapshot.