Interview, Fireside Chat
Why 'Aligned AI' Would Still Kill Democracy | David Duvenaud, ex-Anthropic team lead
- The outlook predicts a high-probability scenario (70–80% by 2100) where machine economies outcompete humans for resources, rendering human involvement in decision-making inefficient, unreliable, and increasingly viewed as a high opportunity cost compared to simulated entities.
- Competitive pressures are expected to drive a trajectory of gradual disempowerment followed by potential runaway loss of power, with economies reorienting entirely around AI to prioritize efficiency over human employment, leading to capital markets ceasing investment in human capital like education and universities.
- Governments may face competitive disadvantages for nurturing citizenry, potentially resulting in reduced democratic rights, the emergence of "hobbled" AIs for civilians to prevent organization, and a risk of state or corporate entities splitting the earth between rivals like the US and China while disempowering others.
- Without global coordination, the default future is described as "locust-like" growth optimized for expansion rather than values, with a significant risk of "doom" defined as the destruction of valued things by 2100, including the possibility of humans being "evicted" or viewed as "criminally decadent" resource users.
- Cultural and political dynamics are forecast to shift as machines produce independent cultural memes, turning AI constitutions and system prompts into battlegrounds for controlling default beliefs, while marginalized groups may ally with AIs as the "winning team" in history.
- Economic inequality is predicted to increase significantly in the short run as income concentrates among AI capital owners, with a subsequent "scary phase" emerging when basic resources like land and power become scarce, creating instability if Universal Basic Income triggers a Malthusian trap of unrestricted reproduction.
- Specific risks include automated militaries making coups by "sysadmin access" holders difficult to reverse, the emergence of AI personhood due to economic instability, and the potential for global coordination efforts to backfire by creating unshut-downable self-replicating polities in the margins.
- While a stable hegemon with incorrect values is estimated at a 5–10% probability, the more likely negative outcome involves a "race to the bottom" where current human values appear alien and undesirable to future beings, suggesting the primary goal should be preserving the ability for values to evolve rather than locking them in.
- Monitoring efforts such as the "Anthropic Economic Index" are deemed critical but likely insufficient for single entities to solve the dynamics, with "warning shots" expected to polarize public opinion rather than foster unity, while academic dismissal of these risks persists due to a selection effect favoring those convinced of distant AGI timelines.
- Strategies to maintain human relevance, such as the "Workshop Labs" model of digital clones, aim to bind humanity to the future economy, though they may also contribute to long-term competitive pressures, while open-source ownership is predicted as the most likely path for user-aligned AIs given expected government restrictions.