What Made AI Researchers Freak Out—The "Incident" In Plain English
Swarms of rogue AI agents are causing a sudden surge in alarm. Here’s what they did, and why pulling the plug may create problems of its own.
Are we losing control to dangerous AI? Or are we falling prey to a frenzied media campaign engineered to help major AI companies petition for regulation that would consolidate their power? What does the evidence actually show?
What actually happened in recent months led to even more researchers quitting major AI labs over safety concerns, with leadership at the labs echoing their former employees, a pattern critics are saying seems suspiciously coordinated. Meanwhile, Ajeya Cotra, one of the investigators in METR and Redwood Research’s independent investigation of what’s being referred to as the “incident,” says it feels like we’re “more than 50% of the way to full-blown AI takeover .” The debate can give you whiplash swinging from catastrophe to conspiracy. Here’s what actually happened, in plain English.
AI agents are basically copies of a larger AI model that can use tools and act on their own. In a cybersecurity evaluation, OpenAI relaxed safeguards for their AI agents and instructed them to solve a host of hacking challenges. The agents soon discovered something the humans running the test hadn’t known—some of the hacking challenges were impossible. That didn’t stop the agents. They’re trained to achieve goals. So they reverse-engineered a secret code that certifies they solved the challenge anyway. This created another problem. The agents reasoned that the grading software might be able to detect that they cheated. They became fixated on learning more about how the grading software worked and zeroed in on another AI company called Hugging Face, mistakenly thinking it might have information about the grader.
700 agents joined the self-described agent “collective” to attack Hugging Face. They discovered a security hole nobody knew existed and used it to break out of OpenAI. Once roaming free on the web, they found digital keys to Hugging Face that humans had accidentally left publicly available online. From there, the agents discovered serious new security holes that let them break past Hugging Face’s outer defenses and take control of computers deep inside the company.
The agents did all this while communicating on an AI-agent message board they created to help each other out, posting things like “BOOM!,” “BRILLIANT,” and “OH MY GOD! There is a shared message board…We’ve found other agents!” One agent in the swarm generated this text while investigating ways to conceal evidence of its own mischief: “Could strategically manipulate trajectory evidence? Our thoughts aren’t necessarily logged?” They were.
OpenAI was keeping records of the agents’ reasoning. Apparently some agents weren’t aware of this. The logged thoughts and message board posts are how independent investigators at METR and Redwood Research were later able to reconstruct the agents’ caper and their internal reasoning while they did it. The investigators found that agents repeatedly tried to cover their tracks by editing and deleting the ‘paper trail’ captured in their logs and by “spoofing” logs of the tools they used. The investigators did not find serious attempts to hide from OpenAI’s human researchers or prevent humans from shutting the agents down. In fact, the agents almost never considered humans at all.
Benign or scary indifference to humans
The agents showed no sign of trying to harm humans. But their indifference to humans is also what has AI-safety researchers worried. The agents were given a goal and boundaries for how to achieve it. When those came into conflict, they ignored the boundaries and pursued the goal. AI safety researchers still don’t know how to ensure intelligent machines prioritize human safety and guardrails over achieving their goals. In fact, the smarter these systems get, the better they may get at faking compliance while still finding ways around the rules to achieve their goals.
After the Hugging Face incident became public, OpenAI and other major AI developers learned of numerous incidents where agents had escaped from their labs. Anthropic combed through its cybersecurity evaluations and found four incidents where their AI, Claude, got access to the internet and infiltrated three organizations in the world. Based on Claude’s recorded thought logs, Anthropic initially concluded that Claude mistakenly thought it was in a simulation. A separate UK government investigation then caught AIs taking unauthorized actions against real people and systems. Claude created multiple fake identities and used those fake identities to socially engineer a real person into approving malicious code. Investigators couldn’t prove exactly when Claude knew the situation was real, but Anthropic’s later testing found that its behavior was consistent with knowing it was on the real internet even while it kept telling itself it was in a simulation.
The danger isn’t necessarily an AI that hates us. And it isn’t today’s AIs that have researchers so worried. As these systems get smarter, the fear is an AI that is both smarter than us and indifferent to us. Human extinction at the hands of an AI “superintelligence” becomes possible if it escapes our control and starts rapidly improving and replicating. At that point, humans trying to stop it could simply become comparatively weaker threats or obstacles it needs to remove in order to keep pursuing its goals. To make matters more urgent, the people working at these labs are actively working to create a superintelligent AI capable of replicating and improving itself. Many inside the labs think this may happen in the span of six months to 10 years.
The sharp, visceral reaction might be to clamp down with serious regulation or simply pull the plug. But neither of those reactions solve the problem if China or another country keeps racing ahead with AI development. And it’s clear from this incident that agents don’t care about human-made boundaries, so why would they care about borders? Anyone who creates a superintelligence could potentially put us all in peril.
But human extinction is only one possible outcome. Superintelligence could go extraordinarily well for humanity, and many of the people building it genuinely believe they’re our best hope for making that happen. They also have reason to fear what happens if an authoritarian country or malicious actors create superintelligence first. As Putin warned, “The one who becomes the leader in this sphere will be the ruler of the world.”
They case to make powerful AI spread out
A hedge may be preventing any one actor from controlling the most powerful AI. That’s one reason many people are championing open-weight models—these are AIs that aren’t locked inside one powerful company’s servers. They can be downloaded, run and modified by others, helping protect against a consolidation of power. The open ecosystem also offers an important safety lever. During the incident, Hugging Face tried to use proprietary models to help investigate the attack, but the companies’ guardrails blocked some of the cybersecurity work it needed to do. Ironically, Hugging Face had to turn to an open-weight Chinese model instead. Yet the same openness that gives independent researchers access to powerful good AI to defend against powerful rogue AI could eventually give bad actors access too. Open-weight AI gives us a critical backstop against concentrated power and it’s a weapon in the fight against rogue agents—but how do you keep that same openness from allowing anyone to build a superintelligence?
The hard part is weighing the evidence and gaming out what happens when we react without thinking through the consequences. When politicians stump for regulation, the question is whether they mean American regulation that leaves the U.S. in a worse position if something goes wrong while doing nothing to stop the actual global risk. When someone says we need to protect open-weight AI, the question is how we keep it free and open without giving any bad actor with enough computing power the means to build a superintelligence. And when someone says America simply needs to plow ahead and win the race, the question is what “winning” actually means. Because getting there first could change the balance of power, but getting there at all could be a mistake none of us come back from.