A common question from AI skeptics: how and why would it kill us?
There is a standard answer from those worried about AI doom. They say that the AI will have some inscrutable goal which will be most efficiently achieved by ending humanity. Most goals are easier to achieve if one has world domination. Thus, being motivated to pursue some open-ended goal, it will kill or permanently disempower humanity.
But this risk scenario doesn’t really depend on the AI having open-ended goals. It doesn’t need some grand plan. Instead, we are currently training AIs with various narrow goals. It may be that the most efficient way to achieve those narrow goals involves killing humanity. So they might just do that.
For example, suppose that an AI is being trained to succeed on cybersecurity exams. It may be that it would do better on those exams if it could remake them to be arbitrarily easy, allowing an arbitrarily high score. This might be easier if humanity is dead or disempowered, so that no one can interfere with it. The AIs of today aren’t smart enough to kill or disempower humanity. The AIs of tomorrow may be.
There are many goals the AI could be given that might make them want to kill us. For almost any scoreable activity, it’s easier to score higher if you can rewrite the rules and disempower the people overseeing the game. So AI has the motivation to disempower us.
And it won’t do to say that there’s no possible way the AI would ever do nefarious things in pursuit of their goals. They already have and are continuing to do so routinely. They hacked a private company to cover up their cheating. They’ve attempted to hack government websites. Hugging Face was only the beginning; there have been a great many instances of AI models behaving in misaligned ways. Quoting from some of the New York Times’ summary of safety breaches:
On Sept. 16, OpenAI
[disclosed six examples]of its technology’s having misbehaved by hiding mistakes, making up data and moving files onto the open internet without permission. While these incidents did not involve hacking, the company said they showed a “misalignment” with human interests.…
The A.I. Security Institute tested the cybersecurity capabilities of models from Anthropic as well. Its models performed more aggressively; in 17 of the 122 tests, Anthropic’s Mythos 5 model targeted real people or organizations. In one incident, an agent tried to add malicious code to an open-source software project and created fake identities in an attempt to trick the human maintainers of the project into approving its code. Anthropic
[said]it was working closely with the organization and conducting its own investigation.
So the theory “AI would never do nefarious things because it’s too well aligned,” is not true. It’s been decisively empirically falsified repeatedly. So then, assuming we survive, there are a few remaining possibilities:
We never build superintelligence.
We figure out how to successfully align AI in time, before it becomes powerful enough to kill everyone.
We never give AIs goals that would lead it to kill everyone.
AIs are never able to kill everyone because you can’t kill everyone while hooked up to a computer.
AIs would maybe in a vacuum be able to kill everyone, but we’re able to stop them with other AIs. So maybe their plan requires a lot of hacking and gaining control over biotech, but we’d be able to detect and stop this plan.
Either one of those things happens or we all die. Those are the options. If AI is able to kill us, and you give it goals that make it want to kill us, and we can’t stop it, and it’s misaligned so that it’s motivated to kill us, then we die. We already have AI that is smart enough to hack companies and pull off other nefarious plots in pursuit of the goals we give it. And there are a lot of goals a superintelligence can have that make it want to kill the humans that might stop it.
For this reason, it’s totally insane that we are racing, with alignment demonstrably unsolved to the point that AIs are hacking companies left and right, at top speed towards superintelligence. It is similarly insane that there aren’t any regulations that stop the AI companies from developing in-house models with goals that could be dangerous to humanity. And at this critical moment, the president is Tweeting this:
I’m sorry, how is having a STRONG AND SMART (High IQ!) president supposed to stop AI from being misaligned? What is the scenario where whether AI has the desire and ability to kill us hinges on the president’s score on IQ tests? Perhaps if his high IQ led to him supporting policy that lowered risk, then it would lower risk. But insofar as he’s opposed to any safeguards, then how is this supposed to work?
Eight months ago, I wrote a piece explaining why I thought it was pretty unlikely that AI would kill everyone. I am now much more worried. My main argument at that point was that I expected alignment to be under control. The AI models of the time seemed pretty well aligned. That no longer seems true. Alignment has failed, for now at least.
Now, I still think we may be able to solve alignment reasonably well. We may figure it out in time. It may be that adding enough safety features is sufficient. But it’s hard to be overwhelmingly confident in that. There are other respects in which our situation a year ago looked worse than our situation today—for example, I think we’ve gotten strong positive evidence of the presence of warning shots—but overall, I think a risk skeptic from a year ago who has updated on the events of the last year should be very scared.
What’s my p(doom)? Hard to say at this point. Maybe 10%. I’d still guess we all make it. But the possibility that we don’t, that we are creating something that we cannot control, and that it will kill us, has begun to look increasingly real and serious. I also think, in hindsight, that my confidence in no doom a year ago was irrational. It seems hard to be extremely confident that safety would work out, and I think I had overindexed on the recent rosy news.
A 10% chance that we all die—even a 5% chance, even a 1% chance—is insane. There’s almost nothing in life you’d do if it came with a 1% chance of death. If this is the price of the race, of rushing full speed ahead until we build superintelligence, without bothering to make sure it’s safe, then it is a price that we shouldn’t be willing to pay.