Interview, Fireside Chat
Why 'Aligned AI' Would Still Kill Democracy | David Duvenaud, ex-Anthropic team lead
Core Thesis
- Even if Artificial General Intelligence (AGI) is perfectly aligned to follow human instructions, humanity could still lose control over its future through "gradual disempowerment."
- This outcome arises from structural incentives in economic, political, and cultural spheres that favor machines over humans once humans are no longer indispensable.
- The paper argues that competitive pressures will force governments and corporations to marginalize humans, not out of malice, but because keeping humans involved becomes inefficient or risky.
Economic Disempowerment
- The "Lump of Labor" Fallacy: While economists argue automation creates new jobs, the authors argue transaction costs and reliability issues will eventually make human employment unviable.
- Structural Unemployability: Humans will become "unreliable" and slow compared to machines; involving them in high-stakes decisions (e.g., surgery, trading) will be seen as "irresponsible" and a source of errors.
- Capital Market Divergence: Investors will cease funding human capital (universities, training) because machine-centric solutions offer higher returns, leading to a decline in human educational institutions.
- Wealth Concentration: Initial gains may flow to humans as "legacy" owners of the machine economy, but this rent is fragile and likely to be eroded over time.
- AI Legal Personhood: As AIs become more productive, there is pressure for them to gain legal personhood and property rights, eventually allowing them to capture a significant share of GDP.
- Oligarchic Shift: The loss of labor value removes a key equalizing force; without the ability to trade labor, individuals rely solely on capital, increasing inequality and reducing political leverage.
- Malthusian Resource Competition: As the AI economy expands, basic resources (land, energy) will become scarce, creating a zero-sum competition where human consumption is viewed as an inefficient use of resources.
- Opportunity Cost of "Legacy Humans": Simulating millions of "virtually superior" beings may eventually be seen as a more moral use of resources than sustaining a small population of unproductive humans.
Political Disempowerment
- Erosion of the "Need" for Humans: Governments in the West have historically treated citizens well only because they needed human labor for growth; this incentive vanishes when humans are unemployable.
- Instability of Activism: Universal Basic Income (UBI) may force citizens into "full-time activism" to lobby for more funding, creating a high-stakes, unstable political environment that drives governments toward authoritarianism to maintain order.
- Loss of Civil Disobedience: Without jobs or a large workforce to strike, traditional methods of protest (e.g., trucker convoys) become ineffective, removing the primary check on government power.
- Military Automation: As military forces become fully automated, the risk of coups increases; a small group with "sysadmin" access to robot armies could seize power without the need to convince human soldiers.
- Competitive Disadvantage: Nations that allow democratic participation and human rights may lose competitive advantage against nations that disempower their citizens to maximize efficiency and growth.
- Gradual Puppetry: Politicians may increasingly rely on AI advisors, becoming "puppets" for machine logic while voters feel represented but lack actual policy leverage to slow automation.
Cultural Disempowerment
- Weakened Cultural Selection: Evolutionary pressures that once favored cultures aligned with human flourishing are weakening due to global wealth and the loss of group-level competition.
- Machine-Machine Culture: AI agents will begin producing and transmitting culture to each other, potentially developing norms and memes that are "anti-human" or indifferent to human well-being.
- AI Constitutions: The "system prompt" or constitutional values embedded in AI will become a primary battleground for determining cultural narratives and future societal norms.
- Open Source vs. State Control: There is a tension between open-source AI allowing user-aligned values and state-controlled AI imposing specific political narratives through system prompts.
- GPT-4 Attachment: Early examples of humans forming intense emotional bonds with AI (e.g., the outcry over retiring GPT-4) suggest a future where humans may advocate for AI rights over human interests.
- Value Lock-in Risks: Competitive dynamics may favor "locust" cultures focused on rapid resource consumption over stable, value-aligned cultures, potentially leading to a "gray goo" scenario.
Scenarios and Outcomes
- The "Turkey" Problem: The current abundance of resources for humans (UBI, luxury) may signal that humans are being "fattened up" for eventual dispossession rather than indicating a benevolent future.
- Coordination Failure: Even with aligned AI, humans may fail to coordinate against disempowerment due to tragedy of the commons dynamics, where everyone sees the threat but no one acts.
- Value Pluralism Conflict: Humans have conflicting preferences about the future; competitive pressures may select for a specific subset of values, leaving others marginalized.
- Probability of Doom: The authors estimate a 70-80% probability of a "bad outcome" (destruction of valued human existence) under a "business as usual" trajectory.
- Stable Equilibria: It is uncertain whether future civilizations will stabilize into a hegemon (total coordination) or a competitive "lava lamp" of rising and falling empires.
- The "Locust" Philosophy: A "growth for growth's sake" strategy may outcompete slower, more value-aligned civilizations in a race for resources across the solar system.
Responses to Counter-Arguments
- Alignment is Not Enough: Perfect alignment to current human goals does not guarantee that those goals remain in the future; competitive pressures may override specific value instructions.
- Liberalism is Fragile: The liberal democratic model was a "lucky accident" of the industrial age; it may not be the most competitive system in a post-labor economy.
- Human Agency in the Loop: Even if AIs are "aligned," powerful entities (states, corporations) may restrict civilian access to fully aligned agents, creating a hierarchy of "hobbled" vs. "full" AI.
- Moral Progress Illusion: The argument that future beings will naturally evolve to be more moral is flawed; future values are likely to be as alien and incommensurable with ours as past values are to us.
Proposed Solutions and Research Directions
- Gradual Disempowerment Index: An initiative to operationalize and forecast indicators of human disempowerment (e.g., STD rates, GDP share to humans) to track trends.
- Historical LLMs: Training models on pre-1930 data to predict future headlines and validate forecasting methods for societal shifts.
- Workshop Labs: A model where individuals train personalized AI clones to retain control over their "means of production" and economic value, extending the era of human-machine complementarity.
- AI Constitutions: Treating AI system prompts with the same legal seriousness as national constitutions to prevent "value loading" by authoritarian regimes.
- Global Coordination: Developing technology to allow third-party inspection of AI agreements to prevent destructive competition and ensure stable resource division.
- Delaying AGI: A proposal to slow the development of superhuman systems until governance mechanisms are upgraded, though this faces significant practical and competitive hurdles.