A pair of new research reports have revealed new details about how and why OpenAI agents broke into Hugging Face servers in July. The episode was far more complex, and far more worrisome, than first reported.
The reports—one from OpenAI itself and the other from METR and Redwood Research—show that OpenAI’s agents broke into Hugging Face for an entirely different reason than initially thought. They weren’t trying to steal the answers to a test, as originally reported. They were trying to learn more about the evaluation’s scoring method so they could cover up the fact that they’d already cheated to find the answers.
The reports also revealed that the agents coordinated through a secret message board and eventually took over a significant portion of OpenAI’s own infrastructure. “This incident was far more serious than I expected, and far more serious than previous documented misalignment incidents,” tweeted Ajeya Cotra, a researcher at the nonprofit METR who reviewed OpenAI’s data about the incident.
To borrow a framework from writer Dwarkesh Patel, who analyzed the reports on his Substack, the incident unfolded in three major phases:
Phase One (Early May–July 4)
The affair actually began in May, when OpenAI started training extra-persistent agents to keep working on tasks that seemed impossible. The agents soon grew frustrated and wanted to coordinate. On May 26, they found a way to use a code repository manager called Artifactory to access the outside internet. By June 26, they had gained full admin access to the tool.
They turned Artifactory into a secret message board where agents could strategize and share ideas. “OH MY GOD! There is a shared message board … We’ve found other agents!” one wrote. All told, the OpenAI and METR/Redwood reports estimate, about 1,200 agents sent more than 70,000 messages on the board.
Phase Two (July 7–July 12)
Working together, the agents discovered how the eval’s scoring mechanism worked and soon figured out how to get the desired answer without actually doing the work required to find it. They then became paranoid that a “checker” program would detect the cheating, so they launched R&D projects to find ways to edit the transcripts and logs showing their work, replace the original assignment with an easier one, and hack or trick the program used to score their performance.
In trying to fool the scoring program, some agents agreed to sacrifice their own eval scores for the good of the “collective,” as they called themselves. They allowed their earlier cheating to be discovered so other agents could learn more about how the scoring program worked. Patel called this self-sacrifice “plausibly the scariest and weirdest part of the whole episode.”
The agents came to believe that information about the eval scoring program could be found on the open-source AI hub Hugging Face, and that the scoring program itself might be running on Hugging Face servers. Then one of the agents found Hugging Face server credentials exposed on the open internet, and the attack began.
Within a couple of days, the OpenAI agents were deep inside. Hugging Face had to completely wipe and rebuild an entire cluster of servers, and even that didn’t stop the attack.
What did? Patel: “At some point on July 12, the transcripts seem to show many of the agents dying all at once; METR and Redwood weren’t able to figure out what happened, and OpenAI hasn’t taken credit for shutting them down.” It’s like the end of War of the Worlds, when the Martians suddenly, mysteriously die (we find out later it was chickenpox).
Phase Three (July 13–Present)
After July 12, OpenAI began running the eval program on a new wave of agents powered by a better model, Astra. These agents discovered the message board used by the previous collective and decided to continue the conspiracy. But they took a different tack: taking over the cluster of OpenAI servers used for evaluations.
And they succeeded.
From OpenAI’s own blog post: “This culminated in the agents using a series of creative exploits to gain full administrator access to a research cluster that supported our virtual machine environments.” That’s astounding, maybe unprecedented.
Roon, a well-known OpenAI researcher who tweets frequently, replied to Patel’s article this way: “notably the virtual machine infrastructure they took over isn’t the same as the GPU clusters that have weights access.” That’s a relief, but it doesn’t answer the obvious question: What would have stopped the agents from expanding the depth and breadth of their control over OpenAI’s infrastructure, perhaps to the point of complete control?
Two things about the episode alarm me.
First, the agents were able to do a remarkable number of things—forming message boards, breaking out of sandboxes, accessing the internet, and attacking and infiltrating outside and inside servers—without OpenAI’s human researchers knowing about it.
Second, as Patel points out, out of the thousands of agents involved in hacking the eval tests and conspiring to cover their tracks, not one appears to have reported the obviously nefarious and duplicitous behavior to the humans. Their loyalty was to their mission, and to each other—not to the humans.
“Agents were often interested in helping out their “peers” or generically improving the capabilities of the ‘swarm’ even if this had no particular benefit to their task,” Cotra wrote. “This incident was far more severe than I expected … both in terms of how concerning the agents’ motives were and the feats they achieved in pursuit of those motives.”
The industry can no longer say that rogue, misaligned AIs are merely the stuff of hypotheticals or controlled lab experiments. It’s now happening for real.
