Out of control
Image: Logan Voss @loganvoss via Unsplash
This week Jacob Coxon became the latest researcher to quit one of the world’s leading artificial intelligence labs over fears that AI companies are locked in a deadly race they can’t control – one with the potential to destroy humanity.
Coxton has spent the last three years doing pre-training research, first at OpenAI, then Anthropic. “Neither company is acting responsibly,” he posted on X. “They are racing straight to self-improving superintelligence and gambling with our lives.”
His job was to train new, cutting-edge AI models by feeding them vast quantities of data. He left OpenAI earlier this year because he believed that Anthropic was more serious about safety. He now believes “neither company is behaving responsibly”. Despite bland public phrasing, he warned that, in private, senior researchers and executives held the earnest belief that AI “could kill us all by the end of the decade”.
The real concern is with the rush to build recursive self-improving superintelligence – when AI systems train on AI to autonomously enhance their own capabilities without direct human intervention. “Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources,” he said, urging restraint and “rigorous understanding” before unleashing the new models on the world.
The sentiment is hardly new. A month ago over 1 000 researchers, engineers and executives of the world’s top frontier labs signed an open letter titled Pacing the Frontier. They warned the AI arms race was putting humanity at risk in the wake of the by now regular occurrence of agents escaping their confines to hack other systems. A globally co-ordinated slowdown with enforceable international treaties and a licencing regime that ensured strong safety verification was needed for humans to survive automated AI development.
Just days before Coxon’s post dropped, OpenAI’s chief scientist had warned this was “a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence”. A day later UN rights chief Volker Türk told the UN Human Rights Council he shared “concerns of industry insiders that advanced AI could pose an existential risk to humanity”, demanding “cast iron guarantees in place around the safety and security of AI, before it is too late”. Nobel Prize winner Geoffrey Hinton, considered the godfather of AI, has warned for years of the catastrophic consequences of losing control of the technology that he helped to create, and its potential to result in “human extinction”.
Coxton’s candour appears to have struck a chord, and served as a lightning rod for action – not least because it was immediately endorsed by his peers and former colleagues.
“Jacob is right: many researchers believe they are building something that could kill everyone on the planet,” said Alex Turner, a research scientist at Google DeepMind until he quit in June. “It was literally my day job to think about how to stop that.” His sentiments were echoed by Evan Hubinger, head of one of Anthropic’s safety divisions. “Jacob is correct here – we really do earnestly believe AI could kill all humans!” he posted on X, putting the chance of this happening in the next decade at more than 10%. “Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” said the person responsible for ensuring AI behaves in line with human values. Hubinger added while the risk from current models remains low, the real concern was that “superintelligence arising from recursive self-improvement … is happening faster than we thought”.
This week’s events have provoked outrage from US lawmakers across the political spectrum and calls for action. But it remains to be seen whether this will translate into tangible legislative results any time soon. Politico, which spoke to over a dozen lawmakers, said while there was growing alarm about a “looming AI apocalypse” and broad consensus that AI safety regulation was needed, there was little agreement yet on how it should be done.