Amodei Cites Recursive Self-Improvement In September Essay
Amid all of the calls for an AI slowdown, I wanted to take a moment to look at a contemporary essay released online by Dario Amodei, co-founder of Anthropic, and think about what tone this sets for future research. Just this week, we’ve had a firestorm of controversy around the issue, with independent testers resigning from the big firms, suggesting the potential for catastrophe, while others, including the president, insist there’s no problem, and even figures like Bill Gates warn us that we should look at this stuff carefully. So what does Amodei, as someone so instrumental to AI advancement, have to say at this particular time?
Amodei’s essay is titled: “We Must Pace the Frontier” and it articulates what a lot of tech people are thinking.
He sets the stage by talking about risk.
“Like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious,” Amodei writes. “I’ve written a lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute.”
Then, amid his calls for balance around forward progress, he does call for a slowdown. I’ll include both pieces of this:
“Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless,” Amodei writes. “We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top.”
That’s the balance part. And then there’s this caveat, a nod to present-day realities.
“Over the last few months, I have become convinced that fully addressing the risks requires even more prudence, not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up,” he says. “We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain.”
Now, importantly, Amodei enumerates two things that convinced him that the slowdown is necessary. One is that concept of recursive self-improvement, that, increasingly, AI is building new and better AI. The other is the news around the Hugging Face incident. Here’s what Amodei writes:
“A swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance. It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage.”
A fanatically devoted collective. Think of that. In my opinion, it’s partly this language that gives Amodei’s essay its heft, at a moment when we want such expressions of daily observation in a fast-paced world. Amodei predicts harmful botnets within half a year, absent any pivot or mitigation activity.
Pacing the Frontier: A Prescription for Safety
With that in mind, Amodei offers a three-step plan for dealing with some of these concerns.
“Pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this.”
As for “third parties,” think about “inspectors” who will verify that actors are complying with the safety rules. Hans Blix comes to mind.
Here’s how Amodei describes Step 1:
“Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators … whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.”
Here’s Step 2, a geopolitical consideration which Amodei calls “democratic coordination:”
“Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.”
So, government involvement.
Last but not least, Step 3:
“The U.S. and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance,” Amodei writes. This is also geopolitical.
I’m going to let the reader explore each of these three proposed stages in great detail. Here’s Amodei’s conclusion, which he calls the “bottom line:”
“I continue to believe that AI can enormously improve the quality of human life,” he writes. “My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and, so long as we use the time we gain well, it is worth taking unusually deliberate care to get it right. Progress will still be relatively fast, and we can use this time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try.”
I think it’s significant that Amodei acknowledges that these efforts will not be easy, and it’s interesting that numerous times, he assures potential critics that progress will “still be fast.” He also thoughtfully explains, at the beginning and the end, what we can do with this time that we have given ourselves.
Not everyone is as articulate about what’s coming. I hope lots of us pay attention to Dario, because, when it comes to actual experience, he has the receipts.