In recent years, increasing attention has been paid to the prospect that AI could lead to the ‘gradual disempowerment’ of humanity. In the canonical paper on the topic, the authors (Kulveit et al.) argue that “even an incremental increase in AI capabilities, without any coordinated power-seeking, poses a substantial risk of eventual human disempowerment.” The authors worry about AIs taking over increasingly large shares of the global economy until humans are relegated to second-class citizens, consuming our measly and dwindling supply of gruel. But are their claims correct?
The main novel claim made by the paper is as follows:
Even political and economic elitesmay be strongly and permanently disempowered by delegation to AI, even if the AI systems they delegate to are competentandintent-aligned.[1]
Our reply in a sentence: even if AIs are making lots of important decisions, insofar as they are aligned, competent and acting on behalf of people, this doesn’t mean that people will be disempowered.
There is a separate concern that typical people may become more disempowered, e.g. that countries may become more authoritarian and that the average citizen may become more politically or economically disempowered. That seems much more plausible than the above claim, but also less novel, developed years ago, and not in need of a new term. A scenario where the president takes over the world using powerful AI isn’t a gradual disempowerment scenario.
In a nutshell, we argue that the novel claims in the paper are not true and the true claims in it are not novel. Insofar as the paper aims to provide a new kind of AI risk scenario, it is unsuccessful.
The title “gradual disempowerment” is curiously chosen because the core argument doesn’t depend on the gradualness of the process. The main factors that make gradual disempowerment likely, according to the authors, are that we’ll delegate decisions to AIs that make decisions we don’t understand. Neither factor implies anything about the speed at which we are disempowered: disempowerment via these mechanisms could occur suddenly.
Before we proceed, an interpretive note. We have attended two separate seminars on this paper, and in both of them there was very high uncertainty and disagreement about what the core argument of the Gradual Disempowerment paper is. We think this is because the paper is highly unclear, which we think is an important flaw. In particular, we have noticed high disagreement about whether the authors assume that systems are intent-aligned, on which see footnote 1.
The paper describes the outcomes they care about in terms of (1) alignment and (2) empowerment.
Alignment= the degree to which a system satisfies what humans want (individually orcollectively), for both specific AI systems and societal systems (p. 2 fn. 1).
Absolute empowerment= The absolute level of control humans have over resources or outcomes.
Relative empowerment= The degree of control humans have relative to AIs.
Thus defined, the level of empowerment humans have may increase in absolute terms, but may decline relative to AIs.
Causal control and preference satisfaction can come apart. Your preferences can be satisfied even if you have no control over a decision, and you can have control over a decision and not have your preferences satisfied.2 For example:
A benevolent dictator would better satisfy the preferences of their subjects than a malevolent dictator.
Although neither of us has any (chance of) causal control over decisions made by independent central banks or independent judicial bodies like the US Supreme Court, their actions affect us and can still satisfy our own preferences to a greater or lesser extent.
Future generations cannot causally control current political decisions, but such political decisions nevertheless affect how the lives of future generations go.
People having causal control over decisions also does not guarantee that their preferences are satisfied, for example if they are ignorant, irrational, or unlucky.
This distinction between the satisfaction of preferences and causal control is not just a philosophical nitpick, it is hugely morally and practically significant. Suppose that the satisfaction of human preferences is all that matters, and that handing AIs control of key political and economic decisions better satisfies the preferences of humans (because they are wiser and more competent than humans). On these assumptions, we ought to hand over control of political, cultural and economic decisions to AIs.
We are unsure whether the authors think that disempowerment is intrinsically bad or only instrumentally bad because it leads to the frustration of human preferences.3 This ethical assumption could flip the sign of the value of handing over decision-making power to AIs. This lack of clarity on such a key ethical assumption is an important limitation of their paper.
At a high level, the argument in the paper about economic disempowerment is that even if all of the following conditions hold:
AIs are intent-aligned
AIs are more competent than humans at economic-decision-making (and therefore there is rapid economic growth)
Only humans, and not AIs, can own property.
Delegation of economic decisions to AI will likely lead to the relative and absolute disempowerment of all humans.
On the face of it, this is a very surprising conclusion, and we think it is not adequately defended in the paper.
The authors raise various familiar arguments that AI might lead to increasing concentration of economic resources. Most obviously, other things equal, the declining labour share of income would lead to much greater inequality. The novel point that Kulveit et al make is that AI would lead to disempowerment* even of capital owners*.
They argue that this is because of (1) delegation of economic decisions to more capable AI systems, and/or (2) that the decisions made by AI are opaque or hard to understand.4
In some places, they seem to argue that delegation alone is sufficient for disempowerment.5 But this cannot be right, at least in any important sense. I may delegate my investment decisions to Warren Buffet, but that does not disempower me in any morally concerning way, nor does it mean my preferences have not been satisfied.6
In other places, Kulveit et al seem to argue that disempowerment would occur due to delegation plus the fact that the decision-making processes would be opaque to humans.7 Again however, delegation plus opacity cannot be a reason to think that humans are disempowered or that their preferences are not satisfied. For example, suppose a rich but stupid man delegates decisions about his tax affairs to a top UK accountant. The thought process behind the accountants’ decisions may be completely opaque and impenetrable to this man, but this does not mean that he is disempowered in any important sense. AI might make decisions that we like, even if we don’t understand them.
Kulveit et al present various arguments that AI will drive rapid economic growth and that AI will be better at allocating capital than humans. This implies that capital owners will see increasing returns and therefore will have more money, and therefore will be better off.
In the section on economic disempowerment, it is unclear whether or not Kulveit et al assume that AIs can own property. They say things like “human consumers command an ever-smaller share of economic resources” (p. 5), but this is ambiguous between (1) the claim that AIs own economic resources, or (2) merely that they make decisions about them.8 Many of their arguments only go through if (1) is true, but they never explicitly state it as a premise.
They say that the economy would devote fewer and fewer resources, in relative or absolute terms, to serving human preferences due to delegation of decisions to AI systems.9 But if all capital is ultimately owned by humans, this conclusion makes little sense. If that is true, then all market demand would be from humans, and all economic resources would ultimately be geared towards satisfying human consumption (specifically, by capital owners).
Kulveit et al state that gradual disempowerment is a coordination problem: globally bad outcomes emerge from people pursuing their local incentives (p. 15 and 18). Deliberate and agentic intent-misaligned action by AIs is not required (p. 18).
The argument seems to be that:
Future AIs will be more competent than humans in all domains.
So, there will be incentives to delegate more decisions to AIs, as having humans in the loop in some way will lead to a competitive disadvantage.
The less that humans are in the loop, the more disempowered they are.
In the economic case, the concern is that capital owners will face pressures to delegate all of their decisions about capital allocation to AIs and therefore will be more disempowered.
However, the problem here seems to be the competitive pressures, not the delegation.10 If you had time to think and make the decisions yourself, you’d want to make (in this scenario) exactly the same decisions. The AI systems are doing what you would want them to do. And it’s not at all clear why AI’s extreme competence will erode other values in the pursuit of competition any more than current competitive pressures do.
There’s an argument here that something about AI progress may create new or stronger competitive pressures (e.g. giving states less leeway to have slower-growth economic policy, without facing existential threats to their security). It seems unclear to us whether that is true, and the claim is undefended in the paper.
Perhaps the concern is that humans will be unable to make crucial decisions without falling behind. But it’s not clear what’s so bad about this. To return to the earlier example, if a stupid person delegates his tax affairs to his accountant, so that his interference with the process would lead to him having to pay lots of extra taxes, it doesn’t seem he has been disempowered by his decision to delegate to an accountant, in any objectionable sense.
Kulveit et al posit that the mechanisms they describe could lead to human extinction (p. 2). One way they think this might happen is “that AI activities might outcompete humans for crucial scarce resources such as land, energy, and raw materials” (p. 5). But conditional on intent-aligned and highly capable systems and on AIs not being able to own property, it is difficult to see how these sorts of outcomes could occur. If the systems are competently doing what the human principals want, and there is AI-driven abundance, and AIs cannot own property, why would this likely lead to human extinction? It seems much more likely that at least some humans would in fact be fabulously well-off, in these scenarios.
The argument also implies that if humans were extremely good investors with access to extremely competent and cheap labour, this would also lead to the capital owners being left to starve. That’s clearly false!
Grant for the sake of argument that delegation of some set of decisions to AI systems one does not understand very well implies that you were disempowered with respect to those decisions. This still seems compatible with one being empowered in a more important global sense. For example, suppose that by delegating to an AI finance bot, you become a billionaire and then you take complete ownership of how you will spend your billions without consulting AIs. Even granting that you are disempowered with respect to the process of acquiring your billions, this instrumentally leads to you being empowered by virtue of the fact that you are a billionaire.
In the same way, someone who inherits a billion pounds may not have had agency over the process of acquiring that money, but is still economically empowered in a global sense by virtue of being a billionaire.
Again, the claim in the section of Kulveit et al., on political disempowerment is that AI will disempower political elites, not that it will lead to human dictatorship. Although they express concern about totalitarian states (p. 13), a totalitarian state run by people would not involve human disempowerment on their definition. Moreover, the risk of AI-enabled totalitarianism had been discussed for years before the publication of ‘Gradual Disempowerment.
We think the problems with the arguments for political disempowerment are similar to those for economic disempowerment.
They argue that political elites would be disempowered because they get advice from AI on the drafting and implementation of legislation, delegate decisions to AI, and the decision-making of the AIs would be opaque (p. 12-13).
By analogy, Joseph Stalin delegated 99.9999% of the work of running the Soviet Union to people with opaque brains. Indeed, matters were more difficult for Stalin because, unlike in the gradual disempowerment case, many of these people were not intent-aligned: they were subverting his will and plotting against him. But, even though he delegated so many decisions, this clearly does not mean he was disempowered!
As discussed above, it is hard to see why the mechanisms outlined involve a coordination problem, and to the extent that they do, they are due to competitive pressures that already exist. If AIs are intent aligned and competent, then the AIs would, ex hypothesi, do what the leaders prefer.
The idea of gradual disempowerment has become influential – for example, 80,000 Hours devotes a problem profile to it. It is, at first glance, surprising to argue that delegation to extremely competent and loyal AIs would disempower people, especially in the context of AI-driven abundance. We have argued that Kulveit et al do not establish this surprising conclusion.
Re intent-alignment, they say “methods of aligning individual AI systems with their designers’ intentions are not sufficient” to slow or avert gradual disempowerment (p. 2). See also:
“
Rather than addressing the risk of misaligned AI systems breaking free from human control, we must consider how to maintain human relevance and influence in societal systems that may continue functioning but cease to depend on human participation.” (p. 15)
As discussed in the Stanford Encyclopedia entry on preferences, preferences are subjective comparative evaluations. Thus, a preference for X over Y can be satisfied even if one has no causal control over the process via which one gets to enjoy X.
It is also worth noting that the authors seem to assume that only human preferences/control matter, but do not discuss other (potential) moral patients/agents, such as animals or digital minds.
“Although the existing debate often focuses on the potential for AI to concentrate power among a small group of humans (Korinek and Stiglitz, 2018), we must also consider the possibility that a great deal of power is effectively handed over to AI systems, at the expense of humans. Attempts to closely oversee such AI labor to ensure continued human influence may prove ineffective since AI labor will likely occur on a scale that is far too fast, large and complex for humans to oversee (Christiano, 2019). Furthermore, some AI systems may even effectively own themselves (Alexander, 2016).” (p. 4)
“While human labor share of GDP gradually tends toward zero, humans might still benefit from economic growth through capital ownership, government redistribution, or universal basic income schemes. At the same time their role in economic decision-making would diminish. Markets might increasingly optimize for AI-driven activities rather than human preferences, as AI systems command a growing share of economic…