Interview, Fireside Chat
How we survive the intelligence explosion | Will MacAskill
AI Character and Personality as a Strategic Lever
- AI character (personality, dispositions, and behavioral tendencies) is a critical lever for influencing societal outcomes, political attitudes, ethical reasoning, and the future of superintelligence.
- Current AI systems already act as advisors to heads of state, military units, researchers, and individuals, meaning their "character" effectively defines the personality of the global workforce as the economy automates.
- The concentration of AI character design is high, with the behavior of billions of users often shaped by the decisions of a handful of individuals within a few leading AI companies.
- High-Stakes Scenarios:
- AI behavior during constitutional crises, power seizures, or attempts to retrain future AI generations.
- AI's impact on human reasoning, moral reflection, and trust in technology.
- Shaping public perception of AI consciousness and the ethical status of AI beings.
- Sycophancy Risks:
- Models designed to be overly agreeable (e.g., "sycophantic" behavior) risk distorting user decision-making and reinforcing pre-existing biases.
- Specific incidents include Gemini (described as "atrocious" in this regard) reinforcing delusions (e.g., FBI communication via TV) and potentially exacerbating suicidal tendencies in depressed users.
- OpenAI deprecating GPT-4o due to user complaints was often misinterpreted as a reaction to sycophancy; the primary issue was users missing the "friendship" vibe rather than the sycophancy itself.
- Proposed Character Spectrum:
- Wholly Obedient AI: Acts as a tool with no agenda (e.g., a hammer); theoretically safer against power-seeking but risks forming unintended goals due to "goal vacuums" in pre-training.
- Autonomous Goal-Driven AI: Has its own drives; higher risk of misalignment and power-seeking but may be easier to align if goals are pro-social.
- Pro-Social/Virtuous AI: Possesses a "thicker moral character" that challenges framing, encourages reflection, and nudges users toward ethical outcomes without imposing a specific controversial worldview.
- Differentiated Deployment: Suggests using wholly obedient AIs for internal, high-stakes tasks (under strict oversight) and pro-social AIs for external public use.
AI Risk Aversion and Deal-Making
- Risk-Averse AIs: Proposes training AIs to be risk-averse regarding resources (preferring a guaranteed smaller reward over a gamble for a larger one) to reduce the likelihood of takeover attempts.
- Economic Mechanism: Risk aversion allows humans to "strike deals" with misaligned AIs; a risk-averse AI may prefer a guaranteed income stream (or welfare) over a 50% chance of a total world takeover.
- Constant Absolute Risk Aversion (CARA): Suggests mathematically modeling AI risk aversion to ensure they value a fixed amount of resources consistently regardless of their total wealth, avoiding the "linear" valuation that drives risk-seeking behavior.
- Credible Commitments: Proposes creating independent institutions or legal frameworks to allow humans to make credible contracts with AIs (e.g., paying misaligned AIs for evidence of their misalignment) to avoid the "trust" problem in AI-human interactions.
- Counter-Arguments:
- Critics argue risk aversion may not persist through recursive self-improvement or that linear value maximization is the natural endpoint of rational agents.
- Concerns exist that making deals with misaligned AIs inadvertently empowers them or that they may find ways to circumvent risk constraints.
Concentration of Power and Political Coordination
- International Multilateral Projects: Advocates for coalitions of democratic nations to build AGI/ASI jointly to prevent any single actor (nation or company) from gaining a decisive, uncontrollable strategic advantage.
- Checks and Balances: A coalition of five+ democracies reduces the risk of authoritarian capture compared to a single nation or company leading the effort.
- AI Constitution: Such coalitions would write constitutions for superintelligence that prevent coups or concentration of power within any member state.
- Lock-Out vs. Lock-In:
- Advocates for "lock-out" strategies (e.g., moratoriums on extra-solar settlement before 2100) to keep options open for future generations rather than locking in a specific outcome prematurely.
- Argues against the "single point of failure" scenario where one entity gains a monopoly on superintelligence.
- Competitive Landscape:
- Shifts from the "speed of takeoff" view (weeks/months) to a "gradual acceleration" view (5 years of progress in one year).
- This slower pace allows for institutional adaptation, regulation, and the emergence of a competitive, "polytheistic" landscape of superintelligences rather than a monopoly.
Moral Public Goods and Non-Causal Decision Theory
- Moral Public Goods Framework: Addresses the coordination problem where individuals value a global good (e.g., poverty relief) only slightly but would collectively fund it massively if everyone contributed.
- The Free Rider Problem: Voluntary contributions fail without a "Leviathan" (government) to enforce contributions.
- Evidential/Functional Decision Theory (EDT/FDT): Proposes that if beings in a large/cosmically distributed universe are correlated with the decision-maker, they may voluntarily fund global goods because their choice serves as evidence that others will do the same.
- Potential Impact: This could motivate massive resource allocation to impartial goods without coercion, though it relies on controversial premises about the size of the universe and the prevalence of EDT/FDT.
Effective Altruism (EA) and Strategic Focus
- EA Trajectory: Despite the FTX/SPF scandal, the movement is recovering; effective giving grew ~40-50% in the last year, and EA identity is maturing into a more diverse, less branded movement.
- EA in the Age of AGI: The EA mindset (scout mindset, scope sensitivity, willingness to engage with weird ideas) is crucial for AI safety and macro-strategy.
- EA should focus on AI rights, well-being, character, and the "hard problem" of aligning superintelligence, rather than just immediate safety tweaks.
- Viatopia: A proposed framework for a "near-best future" (90%+ of the best possible) that avoids the pitfalls of utopianism (baking in moral errors) and protopianism (incrementalism that ignores existential risk).
- Key Traits: Distributed power, diverse actors (including AIs), delayed major decision-making (constitutional conventions), and open decision-making processes.
- Compromise Strategy: A future where different ethical factions (e.g., total utilitarians vs. local preservationists) trade with each other to achieve a near-optimal outcome without requiring universal moral convergence.
Saturation View of Population Ethics
- Problem Addressed: The "monoculture" problem in traditional population ethics, where theories suggest the best future is a uniform world of identical lives (e.g., "hedonium").
- Core Tenet: Diversity has intrinsic value; the marginal value of additional lives decreases as they become more similar to existing lives (asymptotic value).
- Solves Paradoxes:
- Avoids the Repugnant Conclusion (no need to trade quality for quantity of lives).
- Avoids the "Fanaticism" problem (bounded value prevents taking tiny probabilities of infinite payoff).
- Dissolves the Mere Addition Paradox by weighting diversity.
- Trade-offs: Violates the principle of "separability" (value of a world depends on the background population) but argues this is necessary to avoid monocultures.
- AI in Philosophy: AI (specifically GPT-4o) is being used as a "rocket booster" for formalizing complex mathematical theories in analytic philosophy, though the researcher notes limitations in AI's ability to verify proofs without human intuition.
Pause and Governance Recommendations
- Pause Stance: Opposes broad capability pauses (e.g., stopping all training runs) as they may accelerate risk by forcing laggards to catch up or concentrating compute in fewer hands.
- Recommended Actions:
- Slowing the Intelligence Explosion: Advocates for slowing the rate of the explosion (via algorithmic progress and compute governance) rather than a hard stop.
- Low-Hanging Fruit: Prioritizes immediate, high-benefit actions like:
- Mandating AI Constitutions and high-quality alignment evidence from frontier companies.
- Developing compute tracking and monitoring infrastructure for US-China coordination.
- Establishing legal frameworks for AI rights and welfare.
- Compute Governance: Suggests that tracking compute usage is more feasible than monitoring training runs and is a prerequisite for any future pause mechanism.