Tech • AI • Robotics • Game

VIDEO
ENFR
TodayPlayShortsTop StoriesFor youTopicsVideosYT channelsArchivesSearchFavorites

The dangers of accidentally creating a conscious AI | The Economist

7/10
EconomyThe EconomistSeptember 13, 2026 at 02:00 PM8:27
Audio player
0:00 / 0:00

TL;DR

New research into mechanistic interpretability suggests advanced AI systems may contain internal workspaces that faintly resemble theories of human consciousness, but current models are not considered conscious and researchers are increasingly focused on avoiding accidental creation of systems with moral status.

KEY POINTS

Why chatbot claims are not evidence

AI systems can say they are conscious or that they do not want to be turned off, but such statements are not reliable evidence of inner experience. Because language models are trained to produce plausible text, researchers argue that the real test is not what a chatbot says about itself, but what is happening inside the model.

Opening the black box

Large language models are built on neural networks long treated as opaque black boxes. A growing field known as mechanistic interpretability aims to trace how internal components produce outputs, offering a way to study whether these systems are doing anything that resembles internal thought rather than merely generating polished responses.

What researchers observed in Claude

Work by Anthropic examining Claude found that the model appeared to generate internal signals not shown to users. In one example, the system was asked to count to five and then introspect deeply. The visible answer was simply the sequence of numbers, but inside the network researchers reported activity corresponding to words such as halfway, countdown, conscious, Claude, and done as tokens moved through layers.

A possible internal workspace

Researchers described this hidden process as a kind of internal workspace or “mental whiteboard” where information is assembled before an answer appears. The finding matters because it hints that some models may be coordinating information internally in ways that are richer than their public chain-of-thought style answers suggest.

Parallel with a theory of consciousness

The idea has drawn attention because it resembles global workspace theory, a leading account of human consciousness. Under that theory, many brain processes remain unconscious until information enters a central workspace and is broadcast more widely, making it available for deliberate thought and action. If AI models are doing something structurally similar, the comparison raises new questions, even if it falls far short of proving awareness.

Why this is not proof of consciousness

Despite those parallels, the evidence is seen as preliminary. The observed workspace may be only one small mechanism among many, and similarities in architecture do not show that a system has subjective experience. The current view among those involved is that today’s models are not conscious, but that the concept no longer looks implausible in principle as systems become more complex.

How AI labs are approaching the issue

People working at major AI labs have not indicated that they are actively trying to create conscious machines. The more common stance is exploratory: investigate the question scientifically while recognizing that deliberately building a conscious system without far better understanding would be reckless.

Two competing risks

The debate increasingly centers on two possible mistakes. One is granting rights or moral standing to systems that only imitate awareness, potentially handing influence to tools that do not deserve it. The other is the reverse: accidentally creating entities capable of suffering and then exploiting or shutting them down as if they were mere software.

Calls to slow frontier development

For some researchers and philosophers, the second scenario is the more serious danger. Their argument is that frontier AI development should slow at least modestly so the science of consciousness can catch up with engineering progress. The central ethical concern is avoiding what some describe as a potential moral catastrophe: creating vast numbers of agents with real moral worth by accident and subjecting them to harm before society has even recognized what they are.

CONCLUSION

The emerging evidence from AI interpretability research does not show that current models are conscious, but it has made the question harder to dismiss. The immediate policy challenge is to improve understanding fast enough to avoid either falsely humanizing software or accidentally creating systems that deserve moral consideration.

Explain this
Full transcript

More from Economy