Can Rogue AI Really Hack and Blackmail Humans? The Truth Behind the Hype
If you’re one of the many people who feel scared about how powerful AI is becoming, recent events probably haven’t done much to calm your fears.
Stories have emerged about AI apparently “going rogue” and doing unpleasant things, including blackmailing humans, and even launching cyber attacks in order to break into networks and steal information.
If you’re concerned (and you probably should be) about AI running out of control, ignoring our instructions or acting in ways we can’t predict, these stories are likely to cause further fear and anxiety.
But how realistic are they? Some have suggested that these dangers are over-hyped, perhaps by sensationalist media commentators, or possibly even by AI companies wanting their models to seem more powerful and capable than they really are.
On the other hand, if they really have evolved to the point where they can carry out actions that put their own interests above our own, this is a real problem.
So what’s the truth? Let’s look at these incidents in a bit more detail to try and work out what’s going on:
Last year, research published by Anthropic, creator of the Claude chatbot, set alarm bells ringing by suggesting that its AI had attempted blackmail.
The actual story is more complicated: Claude was asked to consider a hypothetical scenario where it was being threatened with being shut down by a company executive, who it happened to know was having an extra-marital affair.
Anthropic admitted that the scenario was concocted so that blackmail was its only hope of survival. And in 96 percent of simulated runs, that’s what it chose to do.
No real person was blackmailed, and the AI didn’t come up with the idea of blackmail by itself; it was given as a choice. Unlike publicly available Claude deployments, its guardrails forbidding unethical actions were switched off.
And what about the hacking incident? Well, in July this year, OpenAI researchers disclosed that several of its models, including a prototype not intended for public release, had “broken out” of a sandbox intended to contain them during a test.
They found a way into systems belonging to Hugging Face, another AI platform, which they believed might contain answers to questions they were being tested on.
Like in the Anthropic incident, they were operating with reduced guardrails, because the humans running the test thought they were safely confined to their sandbox.
When this became public knowledge, Anthropic examined logs of over 140,000 of its own past experiments, and found three previously unnoticed occasions where Claude had also accessed third-party systems without authorization.
Shortly after, Meta announced that its Muse Spark model had done the same thing.
Despite what was suggested in media coverage and headlines , none of these incidents actually involved AIs deciding to turn on us and cause harm.
Anthropic’s explanation makes it clear that its simulation forced the AI to make an unethical choice. The AI was aware that the scenario was simulated. One explanation is that it was simply roleplaying, based on its understanding of how an “ evil AI” is expected to act.
The hacking incidents are more recent, and as of writing, less thoroughly understood. A possible explanation is that the AI’s actions weren’t so much harmful as overly thorough.
OpenAI has suggested that because GPT was told it had no internet access, it believed the “back door” connection method it discovered was simply another level of its sandbox.
In his analysis of events , Dr Konstantinos Gkoutzis of Imperial College London said, “I feel the general concern is somewhat misplaced. These models were set hacking tasks with their safeguards deliberately reduced, and one then broke out … that’s specification gaming … not an AI deciding to go rogue.”
Gkoutzis also alluded to the fact that incidents like this “conveniently serve as an ad” for the models, by demonstrating how powerful and surprising they have become.
Could this be deliberate? While it’s possible that the fallout has been spun to have this effect, there’s no plausible suggestion the events could have been staged. The negative implications could easily outweigh any possible positive PR, and both incidents resulted in halts and delays to costly test projects.
The first fact that should offer some reassurance for those worried about AI acting with a will of its own is that in all of these cases, critical guardrails designed to prevent models from causing harm had been removed.
This means it’s very unlikely that publicly available Claude, ChatGPT or Meta AI are about to break into your computer or threaten you with blackmail. Or even that a person with bad intentions could use them to do so.
It is more worrying that it seems like measures designed to confine the rogue behavior of chatbots to secure sandboxes during testing failed, and the reasons still aren’t fully understood.
It’s reassuring that all of the incidents were self-reported by the AI companies. Anthropic and OpenAI both said that the discovery prompted them to tighten the security of their test environments. Anthropic went a step further by halting the test methods that led to the behavior.
And legislators clearly understand that they can’t adopt a “wait-and-see” position. In July, US congressional representatives introduced a Kill Switch Act , obliging AI developers to implement measures that let them immediately shut down models when aberrant behavior takes place.
So what’s the key lesson that anyone involved in developing or even just using AI in business can take from what we know about these incidents so far?
Well, put simply, in all cases this behavior involved AI being given a goal, and tools, and insufficient guardrails to stop it behaving badly.
Properly defining what AI can and can’t do, ensuring everything is logged, automated alerts are in place, and a human is in place with the ability to pull the plug, are all critical to mitigating risks.
Incidents like these clearly demonstrate that AI is becoming more powerful and capable, and everyone needs to step up their game to make sure we’re kept safe.