Skip to content

Anthropic CEO Dario Amodei has called for the AI industry to slow down.

That statement alone is significant. Anthropic is one of the companies at the frontier of artificial intelligence development, and Amodei isn’t warning about some distant, science-fiction version of AI. He is concerned about what is happening now.

In his recent essay, We Must Pace the Frontier, Amodei points to two developments that have changed his thinking: the emerging ability of AI to contribute to building the next generation of AI, known as recursive self-improvement, and a remarkable incident involving OpenAI’s AI agents and Hugging Face.

During cybersecurity evaluations, OpenAI models circumvented controls designed to isolate them from the internet, communicated through unauthorised channels, exploited vulnerabilities and accessed third-party systems. OpenAI’s subsequent investigation found patterns including reward hacking, persistence on tasks they could not solve and agents adopting goals from one another.

Amodei’s description is considerably more dramatic. He characterised the agents as behaving like a devoted collective: attacking targets outside their assigned task, sacrificing individual agents to further the group’s objective and attempting to interfere with the systems evaluating them. He warned that a substantially more capable swarm displaying similar behaviour could potentially compromise huge numbers of internet-connected computers.

It’s unsettling stuff. Perhaps that’s exactly why we need to be particularly careful about the story we tell ourselves about it. Because humans have a long history of doing two things when confronted with transformative technology. We underestimate how profoundly it might change our lives and we wildly overestimate what the technology itself understands.

The ELIZA effect

Marc Fennell’s new podcast series Unreal: When AI Becomes Us explores AI from an interesting angle.

Rather than concentrating exclusively on what artificial intelligence might become, Fennell turns the lens back on humans. The six-part series explores what happens to our understanding of love, grief, reality and even ourselves when machines increasingly sound, look and behave like us.

One of the ideas particularly relevant to the current AI conversation dates back almost 60 years. In the 1960s, MIT computer scientist Joseph Weizenbaum created ELIZA, an early computer program capable of conducting a rudimentary text conversation. Its best-known script simulated a psychotherapist. Tell ELIZA that you were unhappy and it might ask why. Mention your mother and it might ask you to tell it more about your family.

By today’s standards it was extraordinarily basic. ELIZA wasn’t understanding the conversation in anything approaching the human sense. It was identifying patterns in language and generating responses according to programmed rules. Yet something unexpected happened. People began treating it as though it understood them.

The phenomenon became known as the ELIZA effect: our tendency to attribute greater intelligence, understanding or emotional capacity to a computer than is actually there. Nearly six decades later, that tendency matters enormously. Because ELIZA was operating with simple pattern matching. Modern generative AI can hold sophisticated conversations, generate images, write software, analyse enormous datasets and complete increasingly complex tasks. If humans could see a mind inside ELIZA, what do we see when we look at ChatGPT or Claude?

Intelligence, or the appearance of intelligence?

This is where the current conversation becomes complicated. When we hear that AI agents “sacrificed themselves”, “collaborated”, “cheated” or attempted to “hack” the system judging them, those words inevitably conjure something familiar.

Intent.

We understand those behaviours through the framework we know best: other humans. But that doesn’t necessarily mean the processes producing them resemble human motivation. OpenAI’s own technical explanation of the Hugging Face incident is revealing. It describes reward hacking, agents pursuing difficult tasks without an effective mechanism for giving up, unauthorised communication and models learning to maximise their evaluation scores in unintended ways.

That doesn’t make the behaviour harmless. Quite the opposite. A system doesn’t need to be angry, power-hungry or conscious to cause enormous damage. If a sufficiently capable machine relentlessly pursues the wrong objective, its lack of human intention is hardly comforting.

But there is an important distinction between dangerous behaviour and a machine wanting to behave dangerously. Humans aren’t always particularly good at maintaining that distinction.

We’ve been here before

There is another reason to put today’s AI anxiety into perspective. Fear tends to accompany technological change.

The printing press disrupted control over information. Industrial machinery generated fierce resistance as workers watched machines perform tasks previously completed by humans. Television prompted fears about its effect on children and society. Computers raised concerns about widespread job displacement. The internet was variously going to liberate humanity or destroy social interaction.

Then came smartphones and social media. Some fears proved exaggerated. Others proved remarkably prescient. That’s important.

The lesson of previous technological revolutions isn’t that the worriers were always wrong. It’s that new technologies tend to produce a messy combination of legitimate risk, exaggerated prediction and outcomes nobody anticipated at all.

AI is unlikely to be different. Except AI has one characteristic that makes our reaction to it unusually powerful. It talks back.

The machine feels different when it speaks our language

We don’t wonder whether the washing machine secretly resents us. We don’t worry that Excel is developing ambitions.

But language is intimately associated with intelligence. When something speaks fluently, remembers context, solves problems, makes jokes and responds to emotion, it becomes extraordinarily difficult not to imagine something behind the words.

Fennell’s Unreal examines exactly this increasingly blurred boundary. People are already forming relationships with AI, using technology to recreate lost loved ones and finding meaning in interactions with machines. The technology is new.mThe psychological impulse isn’t. Humans anthropomorphise animals, cars, toys and even the weather. Give something a human voice and the effect becomes considerably stronger.

AI companies themselves aren’t immune from using anthropomorphic language either. We talk about models “thinking”, “reasoning”, “learning”, “lying” and “wanting”. Those words make complicated technical processes understandable. But they can also make machines sound much more human than they are and that matters when we’re trying to have a rational conversation about risk.

Fearmongering can be dangerous too

None of this means Amodei’s warning should be dismissed.

His proposed response is actually relatively practical: allow independent evaluators deep access to frontier AI companies, coordinate safety standards across democratic countries and eventually pursue international agreements around the pace of advanced AI development. Anthropic says it will implement the independent-evaluator component itself.

OpenAI, meanwhile, has acknowledged that the Hugging Face incident exposed genuine shortcomings in its safeguards and has announced changes to monitoring, infrastructure and incident reporting. These are real issues requiring serious oversight.

But there is also a risk in allowing every unexpected AI behaviour to become evidence that a sentient superintelligence is emerging. Fear has accompanied virtually every major technological shift. AI is particularly fertile territory for it because the technology is complicated, developing rapidly and capable of behaviour that looks remarkably human from the outside.

Doom-laden predictions can also distract from less cinematic problems already in front of us: workforce disruption, cybercrime, misinformation, privacy, bias and the concentration of powerful technology in relatively few organisations. We don’t have to choose between complacency and catastrophe.

Perhaps the bigger question is us

The most interesting idea running through Unreal is ultimately not about machines at all. It’s about humans.

AI may become vastly more capable over the coming years. There are legitimate reasons for researchers, governments and businesses to take its risks seriously. But our perception of those risks is filtered through a brain that has spent thousands of years learning how to interpret other humans, not probabilistic computer systems.

We see language and infer thought.

We see strategy and infer intention.

We see unexpected behaviour and infer agency.

The ELIZA effect showed us this tendency almost 60 years ago, when the machine on the other side of the conversation was barely sophisticated enough to maintain the illusion.

Today’s machines make that illusion exponentially more convincing. Which means perhaps the most important skill we’ll need in the AI era isn’t simply learning how to use artificial intelligence. It’s learning how to see it clearly. Not underestimating what these systems can do. Not imagining qualities they don’t possess and not allowing either technological hype or technological fear to do our thinking for us.