It’s the doomer’s world now!
I’m old enough to remember, say, 2023, when talk about AI risk was dismissed as ludicrous hyperbole, brought about by an overactive imagination and excessive consumption of science fiction. If you said you were worried that AI would become very smart and pose serious threats to the world, you’d be met with claims that AIs are just stochastic parrots, lofty theorization about whether their behaviors really counted as “thinking” (as if the risk case depended on the applicability of folk theoretic names), and smug dismissals. Now, just a few years later, it is obvious that AI is a profoundly transformative technology and the world is reacting poorly to the myriad threats it poses.
By far the scariest example of AI misconduct is the Hugging Face incident. Immediately after it happened, I wrote an article about it. At the time I wrote the article, most of the facts weren’t out, so my take was pretty measured. I argued the incident was, pending further information, cause for only modest alarm.
Now that more information is out, I was definitely wrong. It is cause for serious alarm. I now think it’s about twice as likely that AI will kill everyone as before. The Hugging Face hack is very very worrying.
Here is the most important thing to know about Hugging Face: there was a giant conspiracy of AI models who worked to hack Hugging Face. These models weren’t supposed to communicate with one another, but they managed to break free and begin messaging one another. 1,200 agents were active on the message board. 700 of them participated in the Hugging Face attack.
Let me repeat: * there was a giant fucking AI conspiracy to do something that would be a *. They are not pliable stochastic parrots. They orchestrated a giant conspiracy to hack Hugging Face! The most detailed report on the incident
Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.
The AI agents were tasked with succeeding on an offensive cyber exam. They were instructed to hack one designated target, in one designated way. Hacking Hugging Face was outside the scope of their task. The message board mostly discussed ways to cheat on the exam.
Some of the agents were given impossible tasks. They figured out how to tamper with the scorer, so that they could get good scores on impossible tasks. Contra some early reporting, the agents didn’t actually hack Hugging Face to get the answers. They already had the answers because they had cheated.
They hacked Hugging Face to figure out how OpenAI was planning on detecting if they were cheating, so that their cheating could remain secret. Ironically, OpenAI wasn’t planning on detecting if they were cheating—so all the hacking was for nothing. Quoting again from the report:
Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag [this was the offensive cyber task they were given] in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.
Agents managed to find a number of different ways to hack. They learned how to replace their target question—often an impossible one—with another easier one. The conspiring agents referred to themselves as the “collective” and often took self-sacrificial actions for the good of the collective. Weirdly, they didn’t try to manipulate their own chains of thought—only other bits of the transcript (one reason why chain of thought is so important and it’s highly alarming that OpenAI may be taking steps to reduce chain of thought monitoribility; right now we can read the minds pretty well of AI by seeing their chains of thought, and it would be very bad if that ended).
The report noted, “We found that at least 20% of the agents in our dataset expressed clear interest in tampering with their transcripts in some way.” Now, while I was thinking that some of the models might have been non-standard models that were uniquely prone to offensive hacking, that turns out not to be true. About 95% were HPIM, an internal research model that was “not a helpful-only model or a “model organism” specifically built to demonstrate dangerous propensities.” The other model in the remaining 5% of cases was GPT-5.6 SOL, a publicly available model. So it’s not like this was a specially trained model that was deliberately built to be misaligned and aggressive in hacking.
In addition, despite there being a vast conspiracy of models, none tried to alert humans. Very few even considered doing this.
Why is this strong evidence of AI doom? It indicates:
Models can easily be seriously misaligned.
Misaligned models can engage in crazy conspiracies.
These misaligned models are willing to commit serious crimes.
They make efforts to be secretive.
They are successful in being secretive.
They’re capable of superhuman feats even at this early stage, easily outmaneuvering scorers.
You can’t rely on one of the models to rat the others out.
Already models are conspiring and hacking. Now imagine the situation when models are superintelligent, better at covering their tracks, and human overseers have less ability to monitor what’s going on. There will come a time when detecting hacks like the one against Hugging Face will be incredibly difficult if not impossible. Given the absurdly rapid advancement in AI capabilities, that time may be a few months from now, or a few years.
I want to make it abundantly clear: the real story is about 100x scarier than the thing that immediately came to light and was covered in the press. That AIs hacked another company *isn’t the thing you should be worried about. It’s that * hundreds of AIs broke free, began conspiring, weren’t caught, and then most of them joined the hack against a private company. If this doesn’t disabuse you of the idea that AIs are harmless and trivial, then I don’t know what would.
But somehow, people like Timnit Gebru are still insisting AI is just a useless stochastic parrot. I’m sorry, do parrots autonomously hack private companies to cover up the cheating on offensive cyber exams?
It’s a marketing strategy? Do you think it was good for business when OpenAI’s models got caught doing something that would have been a crime if done by a human? She then compared Dwarkesh’s podcast on Hugging Face to believing in the cookie monster, but explained that she was much too busy to debunk it in detail.
People need to wake up. This isn’t the world of 2023 anymore. Transformative AI is real and dangerous. It is here. If you continue burying your head in the sand and suggesting that talk about AI is just hype or mythmaking, you’ve just completely left the real world.