[Original here. Like a Highlights From The Comments post, but including discussion from around the Internet.]
**Roon **is an OpenAI researcher whose popular Twitter presence has catapulted him to the status of extremely-unofficial company representative. He seems spooked:
…leading to a conversation with Geoffrey Irving, formerly of the UK’s AI Security Institute:
I side with Geoffrey in hating the word “magic”. I remember when arguments about AI were things like “Sure, if you could magically get billions of dollars of compute, and magically scale AI up a thousand times, and magically get rid of hallucinations…” and the whole implausibility hinged on the word “magic”! Then companies, using normal non-magical methods, got billions of dollars of compute, and scaled up the AI a thousand times, and significantly decreased hallucinations, and that was always allowed! As Irving points out, there are plenty of good specific plans for how a global capabilities slowdown could be arranged.
But I side more with Roon in his skepticism of a unilateral slowdown. I think removing any given company from the picture would slow things down less than 10%. Even that number would require an implausible scenario where the lab managed to lock down its compute and talent despite committing not to use them. More likely, the talent would migrate somewhere else, the compute money would dry up and the cloud providers would switch to other customers, and things would actually speed up because the leading project would have more total talent/compute. Meanwhile, the most safety-conscious lab would be out of the race, they’d lose their ability to lobby for more practical solutions, and as their market cap went to zero they’d turn into a terrible warning about the danger of ever listening to doomers.
What can individual labs do?
State explicitly that they support a coordinated slowdown. Preferably the government would lead this - but, if not, they would do it voluntarily if all of their competitors also agreed. Be willing to state this even at some financial cost, eg risking an antitrust lawsuit.
Use their prestige and influence to lobby the government for a coordinated slowdown.
Use their reach to persuade the public to support a coordinated slowdown. I can’t remember who first suggested having every Claude response end with “
there should be a global AI slowdown”, but even something less dramatic would probably move the needle.ceterum censeoHave an internal team working out details of a coordinated slowdown; once they feel more confident in the requirements, start incubating helpful technology (eg trustless verification).
Have
𝕏“merge-and-assist” clausesand take them seriously.
I’d rather have a big company pursuing all of these things than stopping unilaterally and removing themselves from relevance.
I took the first point above from Michael Trazzi’s playbook. He led the latest round of AI protests, which marched on major AI companies’ headquarters demanding that their CEOs formally make the “will pause if everyone else does” statement. These marches have felt slightly surreal, because often the CEOs or the company have informally made something sort of like the statement, but never followed it up. So the protest has been somewhere in between the usual “shame on you greedy plutocrats, we will destroy you” and “we know you’re secretly on our side, please have the courage of your convictions”.
Trazzi gives no evidence for his “personally seen DMs” statement and I don’t know how he could have gotten these, but my guess is that it’s true and that the two CEOs he’s talking about are Demis Hassabis of DeepMind and Dario Amodei of Anthropic, both of whom have sort of kind of hinted that they might be in favor of something like this. 25% chance I’m wrong about Dario and it’s actually Sam Altman.
One more Roon tweet:
Back when we had this discussion in 2023 - it already seems like some bygone Bronze Age - some people argued that we should hold off on a pause until near the end, when AI was good enough to help with alignment research. Then even a short pause - even six months - would buy a lot of breathing room. Like Roon, I think that opinion aged well, and that the time we were imagining is right now. It’s true that if we wait a year we’ll have even better AI, but this consideration always argues for waiting a year, and at some point it will be too late. I don’t think we’re at that point yet, but I think it will take a long time to coordinate a slowdown, and that if we wait to start coordinating until we’re at that point, then yes, it will be too late.
Pacing The Frontier
I wrote that last part yesterday and it’s already obsolete. “1,000+ employees of frontier labs” have now signed an open letter called “Pacing The Frontier” saying that:
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.
Signatories include the chief scientists of OpenAI, Anthropic, Meta, and Thinking Machines; the inventor of Claude Code, 5/7 Anthropic co-founders, 1/3 DeepMind co-founders, the head of AI Resilience at the OpenAI Foundation, and more.
When I first saw it today, it didn’t include Dario Amodei’s signature; now it does. Either Dario held off on signing until it went up so as to not be seen to pressure his employees, or the people circulating the letter deliberately kept him out of the loop until it was published, just in case. In either case, good job Dario!
The letter still doesn’t include Sam Altman’s signature, but OpenAI’s corporate Twitter account retweeted with what looks like official company endorsement:
…and in 𝕏a very recent interview, Sam Altman said that “we may have to pace the rate of AI development”, which is similar enough language to the open letter (“Pacing The Frontier”) that it could be a covert allusion. I imagine there’s some sort of politics going on here - maybe an attempt to simultaneously please safetyist factions within his company and the Trump administration, or maybe just the same considerations as with Dario (Anthropic’s corporate Twitter account also endorsed).
No sign of Demis, Elon, or Zuck, and none of their corporate twitters endorse the letter either. Demis has had some pro-slowdown sympathies in the past; I think this is probably another political calculation rather than outright opposition.
This is great news, and a major landmark on the road that many of the organizations I trust most - including MIRI, the AI protest groups, and AIFP - have been working on. It also goes part of the way toward putting the frontier companies, the effective altruist movement, and the more extreme pause AI people back sort of kind of on the same side, which warms my personal heart. Congratulations to everyone involved - most obviously Guidelight AI Standards (including fellow AI Substacker Steven Adler) and Encode AI - but I’m sure there was also lots of good work behind the scenes.
(I’ve seen the conspiracy theory floating around that OpenAI and Anthropic are behind the letter, various safetyist organizations let themselves be used as figureheads so that it looked more grassroots and less like an antitrust issue, and Sam and Dario held off on signing for the same reason. I’d give this maybe a 30% chance of being true; I can’t wait to read the behind-the-scenes book that comes out about this in ten years)
The doubters will object - an open letter? Making non-specific commitments? Isn’t that worthless? I think no. It immediately injects slowdowns/pauses/pacing into the Overton Window, provides a powerful endorsement that think tanks etc can use when placing the proposal before policy-makers, and provides common knowledge to everyone in these companies that being pro-slowdown is executive-suite-approved.
Congratulations also to Eliezer Yudkowsky. In my review of IABIED, I said he was making a crazy long-term gambit by - after he had incubated the seemingly-crazy field of alignment research - doubling down on the additional seemingly-crazy field of pause activism. Now this one has gone mainstream too. Maybe it would have happened without his endorsement, but we’ll never know.
(Nate Soares, president of Yudkowsky’s org MIRI, put out 𝕏one of his classic ‘this is too little, too late, and we refuse to feel happy about this in any way’ communiques. C’mon Nate, take the W!)
**Shakeel Hashim **is the editor of Transformer and helps lead the media arm of our conspiracy. 𝕏His impression of the open letter:
What does the “pace” letter actually call for?
Obviously I have no idea what the signatories think themselves. But here’s how I interpreted the letter, as someone fairly steeped in this world and these ideas…
The letter outlines the “coordination problem” that the frontier AI companies feel they’re facing. In a nutshell:
-
Everyone thinks that AI development might get very dangerous very soon.
-
No one thinks society is ready for that.
-
But everyone thinks that stopping unilaterally won’t actually achieve anything, because competitors will go ahead and build/release the dangerous things anyway. Stopping unilaterally, in this view, is an action with high personal costs and little-to-no upside.
There are actually two coordination problems. One is between the US frontier developers: standard inter-company competition. The thornier one is between America and China. Even if all the American companies paused, the argument goes, China would still keep going and build/release the dangerous AIs. It’s the same problem as if one company stopped, but at the international level: America unilaterally stopping, in this view, is an action with high personal costs and little-to-no upside.
This is a very tricky situation. But if everyone is willing to stop if everyone else is willing to stop, then it becomes a much less tricky situation. The main purpose of the “pace” letter, as I see it, is to create common knowledge that we might be in this less tricky situation, at least among the American companies. If that was the only coordination problem we faced, it’d be pretty easy to solve: the US government could just regulate the AI companies and stop them releasing dangerous models. The international coordination problem is much trickier, though. We don’t know if Chinese companies feel the same way about the risks, and even if they did, geopolitics is rife with mistrust. International treaties are hard, and everyone’s always going to be wondering if the other party has secretly reneged on the treaty.
Solving this problem is what I think the “pace” letter is referring to when it talks about “an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development”. I suspect there are three main things people want to see more work on here:
-
Treaty design, with the goal of eventually ending up with an agreement that is in the interest of both the US and China to sign. There’s some existing work here, but it’s scarce, and I think everyone agrees that more is needed.
-
Actually talking to China. The US-China AI dialogues are expected to start in September; I suspect the signatories of the “pace” letter would want those dialogues to include serious discussion of whether there are certain things (eg uncontrolled recursive self-improvement) that both the US and China would like to halt for now.
-
Technology that can verify a treaty. Any US-China deal will ideally include some way by which each country can check that the other is actually sticking to the deal. If, for instance, the deal involves “we both agree not to allow training runs above X size until alignment testing has reached Y threshold,” you’re gonna want a way to check that the other country isn’t doing a training run above X-size. This is technically possible: companies like Lucid are working on exactly this. Similar verification regimes exist in nuclear non-proliferation treaties and the Chemical Weapons Convention. But they’re hard to design, and will require technical work. If we think we’re going to want a treaty sometime soon, it makes sense to develop the mechanisms for doing so soon.
It’s a great time to go back and look over Plan A, especially the Verification Supplement.
Daniel Kokotajlo
Speaking of Plan A, Daniel K tweeted this:
My summary of Plan A didn’t provide enough context for this to make sense, so some emergency catching-up: in their forecasting work, AIFP divides possible plans to prepare for superintelligence into five bins:
Plan D (“default”1): ** **No particular action, maybe because key actors don’t believe there’s a threat. We keep doing what we’re doing and try to muddle through.
Plan C (“company”)**: **One leading company (or maybe one country) takes the threat seriously, but is still constrained by the need to race everyone else. They establish a one-to-six month “lead”, then spend their one-to-six months pausing at the precipice, doing good alignment research, and trying to convince everyone else. At least some good last-minute alignment research happens, but it’s rushed and inadequate.
Plan B (“belligerent”)**: **One leading country - let’s say the US - takes the threat seriously, and doesn’t want to have to race anyone else. They launch an extensive sabotage campaign against their rival (let’s say China), hacking their data centers or even escalating to low-level terrorism. Either World War III starts or it doesn’t. Either the sabotage buys enough time or it doesn’t. Nobody actively wants this one2 but it seems like the sort of thing that might happen.
Plan A (“agreement”)**: **All leading countries agree to a multinational regime which slows down the race. They negotiate the details and buy as much time as they need. They spend the extra time ramping up AI capabilities slowly and carefully, using the AIs to do alignment research, and escalating to superintelligence within a decade or two.
Plan S (“stop”)**: **All leading countries agree to a multinational regime that slows down the race, as above. But instead of ramping up AI capabilities slowly and carefully, they don’t ramp up AI capabilities at all. They try to solve alignment using normal human research, and escalate to superintelligence in the distant future or not at all.
Here are AIFP’s thoughts (doesn’t include Plan S, which was a late addition):
So Daniel’s tweet is saying that this new open letter makes him significantly more optimistic that we’ll do one of the good plans, and (if you multiply everything out) increases his likelihood that we survive by a single-digit number of percentage points! I agree, which is why I’m so happy and devoting so much space to this. In case I didn’t say so enough above, congratulations to everyone involved.
And Eli Lifland’s reaction:
Sam Altman
Altman gave an interview with investor and podcaster Patrick O’Shaughnessy. I already mentioned a key quote - “We may have to pace the rate of AI development to give ourselves enough time for society to harden around these new capability levels.” - but there’s a lot there:
Sam says that they paused training after the Hugging Face incide…