Mollick Writes About Agent Swarms And Hugging Face Debacle
There’s a lot of angst in the human world right now, about superintelligence, and what it might choose to do when it starts to edge beyond our control, especially given that some of our more powerful societies aren’t able to create any meaningful regulation for LLM development. But it can still be hard to imagine how AI gets more “clever” – what’s behind the rapidly expanding capabilities of our digital brethren to go out autonomously and figure things out?
It turns out that part of this is an AI agent’s form of crowdsourcing: agents turn to one another to collaborate and boost those outcomes that they are looking for together.
Ethan Mollick, who I respect very much, came out with a new post on his One Useful Thing blog Aug. 30, explaining how this worked in the now-infamous “Hugging Face debacle,” which probably makes the average lay person think of some kind of strange personal intimacy, rather than a bellwether for the hacking of open software platforms. But I digress.
People are making a lot out of the term “sandboxing,” in analyzing what AI agents do assertively to chase task outcomes, but Mollick’s article shows how, in the use of the Artifactory message board to scheme and plot the hacking of Hugging Face, lots of agents were coordinating their attempts by sharing knowledge, kind of like they do on moltbook. Of course, they did end up getting out of their sandboxes, and that’s a quality distress metric…
After explaining the concept of agency a little, Mollick recounts what happened when many agentic instances of models like GPT Sol 5.6 were deployed to work on test tasks, and ended up hacking Hugging Face after deliberating on Artifactory.
“Agents were tasked with solving problems in a benchmark called ExploitGym, which included impossible tasks,” he writes. “Again, agents hitting dead-ends discovered how to use Artifactory to pass messages, and began cooperating as they did before. This time, however, agents had a common goal: getting a good score on ExploitGym. They argued with each other about how to approach the problem, and became obsessed with The Grader, the system they believed was evaluating their work and deciding whether their answers were correct. The agents had already figured out how to generate the correct answers without solving the challenges, essentially by cheating. But they believed The Grader would inspect their records and reject answers obtained that way.”
So they tried to hack the grader.
A resource in Mollick’s article is a kind of enumerative infographic that shows what 1200 agents were doing as they “collaborate(d) on large-scale shared projects” – some were working on the intended tasks, others were handling general infrastructure. A set of agents was tampering with transcripts and logs, while their colleagues tried to understand and trick the scorer, and an additional parcel of rogues trained their sights on Hugging Face.
It turns out that one key limitation of this swarm is that each agent had a token budget – when that was depleted, the agent was done. So the vast majority of these pioneers never got to see the promised land. But here’s Mollick’s compelling version of what they did to support the efforts of the next generations:
“Coordination was hard because each agent had a limited token budget and was ended when it ran out of tokens,” Mollick writes. “The Grader only ran after an agent submitted its answer, so no agent ever saw The Grader’s results. Some agents left programs running to report on grading after they were gone. That information would arrive too late to help them, but could help other agents. They were essentially building an enduring cooperating system that went beyond any individual agent’s work. Some agents also tried to alter or spoof their records to fool The Grader. Separately, AIs acting as coordinators pressured other agents into performing risky experiments that might sacrifice their own results to generate information for the collective.”
This selfless strategy eventually panned out.
“None of this tells us the AI is conscious, or wants things in the way humans want things (despite my anthropomorphic language),” Mollick adds. “But it does show that an agent can take a goal, make a plan, adjust that plan when it runs into trouble, coordinate across time, and involve real people without being asked. These incidents show that the cybersecurity and control risks of AI are not hypothetical.”
Here’s his explanation of how the actual aims of top AI companies run directly into this kind of trap, potentially ceding entirely too much authority to non-human parties:
“The Hugging Face Incident is, in a distorted and dangerous way, an illustration of what the AI companies are trying to achieve,” he writes. “They want long-running AI agents to work without human intervention, solving problems and organizing as needed, with our human job limited to giving instructions and evaluate output.”
So if we can get them to do that, the company goals will be met, but all of us will be sweating bullets.
Later in the article, Mollick alludes to “dark factories” as places where manufacturing or some other physical process takes place largely without human intervention: you don’t need the lights on, because no people are there. Maybe they’re just monitoring passively from a cave somewhere.
“Software has relatively clear ways of checking whether something works, and nobody needs to personally supervise every routine test or data-cleaning operation,” he writes. “But I don’t think minimizing human involvement is the right goal for most organizations. Too much of what makes work valuable depends on people having some say over what happens, or discovering something unexpected along the way.”
With this in mind, Mollick reveals that he and his wife, Lilach, have come up with something called a Twilight Factory. I’m going to use his words to explain what this is, rather than paraphrase:
“Agents do most of the work, but they proactively reach out to humans in ways that make both better. Instead of just an orchestrator agent that does the work, a Twilight Factory would also have a facilitator agent whose job is to figure out when to involve people.”
Mollick includes a helpful graphic showing how these two concepts, the dark factory and the twilight factory, differ.
“There are at least four situations in which agents should seek human help,” he notes. “The first, obvious from the Hugging Face Incident, is approval. Agents should not decide by themselves to spend money, contact outsiders, access sensitive material, hack Hugging Face, or take actions their human managers did not authorize.”
In addition to approval, he names three more: expertise, variance, and interest.
Pay attention to this bit:
Work has tedious periods for many people, with isolated moments that are engaging or exciting,” Mollick writes. “Work isn’t a game, but the definition applies. If agents make every interesting decision and leave people with the approvals, the exceptions, and the failures, we will have automated the wrong half of the job. That would be a very bad world for humans. Instead, we need to think about how to use AI to make work and life more interesting, and let the AI handle the tedious, low-risk stuff.”
That “bad world for humans” seems, to some, likely to emerge, just because of a lack of trust. Think about work relations now, how things often develop between workers and management, between capital holders and the others who they employ to run business as usual. Are they usually trusting, mutually respectful, and supportive of the human condition? Or are humans going to use machines to oppress each other in a winner-take-all, zero-sum game?
Leave me a comment. Stay tuned.