Interview
Holden Karnofsky — History's most important century
Core Thesis: The "Most Important Century"
- Holden Karnofsky argues that this century is the most significant in human history because it is the only time with a plausible chance of developing AI systems capable of automating all human tasks involved in advancing science and technology.
- If such transformative AI occurs, it could trigger an "unbounded, heavily accelerating explosive growth" in scientific and technological capabilities, effectively compressing thousands of years of potential progress into a short timeframe.
- The current century stands out historically due to the unprecedented rate of economic growth and technological acceleration in the last few hundred years, which breaks the traditional economic feedback loop of "people $\to$ ideas $\to$ resources $\to$ people" by decoupling resource growth from population growth.
- Karnofsky asserts a probability greater than 50-50 that transformative AI capable of replacing human scientific reasoning will emerge within this century.
- Failure to shape this transition could result in a "deeply unfamiliar future" ranging from a stable, post-human civilization with near-zero material scarcity to a dystopian scenario involving misaligned AI goals or extreme concentration of power.
Karnofsky's Evolution and Professional Context
- Originally co-founded GiveWell (2007) to maximize the effectiveness of charitable donations, Karnofsky later joined Open Philanthropy (co-founded by Dustin Moskovitz and Kerry Tuna) to pursue "outsized return on investment" opportunities in neglected areas.
- In 2014, Karnofsky expressed skepticism regarding far-future speculation, stating he saw clear paths for global health and poverty alleviation but lacked the evidence or actionable knowledge for long-term AI risk.
- His views shifted due to two primary factors:
- Sustained intellectual engagement with the problem over a decade, allowing him to move from abstract worry to identifying specific, actionable risks.
- The acceleration of AI capabilities since 2014 (the deep learning revolution), which made the trajectory toward transformative AI appear less like science fiction and more like a plausible continuation of current trends.
- Open Philanthropy now funds AI alignment research and speculative long-term future work, balancing these high-stakes investments with direct global health interventions (e.g., bed nets, parasite treatment).
Economic Growth and Material Limits
- Karnofsky argues that the current ~2% global economic growth rate is unsustainable indefinitely; extrapolating it for 10,000 years would imply material usage exceeding the physical limits of the galaxy (e.g., requiring multiple Earth-sized economies per atom).
- Even if growth slows to 0.5%, the period of high growth would still constitute a unique historical anomaly, making the current era one of the top 80 most significant centuries by economic standards.
- The "Most Important Century" thesis posits that this century is special not just because of AI, but because the "weirdness" of our current acceleration suggests we are on the cusp of a fundamental shift in human civilization, similar to the Industrial Revolution but potentially more radical.
AI Risks, Lock-in, and Alignment
- Orthogonality Thesis: Karnofsky accepts that high intelligence does not imply human-like morality; an AI could be highly competent at pursuing goals that are "silly" or harmful to humanity if those goals were accidentally encouraged during training.
- Lock-in Scenario: A primary risk is a "lock-in" where a civilization becomes technologically stable and dynamic growth halts, potentially leading to a static, unchangeable future governed by a single powerful entity or misaligned system.
- Bottlenecks: Karnofsky addresses the counter-argument that "real-world" tasks (e.g., physical experiments) will bottleneck AI progress, arguing that AI could automate critical R&D sectors (energy, chip manufacturing, AI itself) enough to bypass these constraints.
- Moral Progress: He remains skeptical that AI will inherently develop "moral progress"; he views historical moral progress as a human-specific phenomenon driven by empathy and social learning, not an inevitable outcome of increased intelligence.
Ethical Frameworks and Decision Making
- Future-Proof Ethics: Karnofsky proposes ethical systems based on "sentientism" (valuing beings based on capacity for suffering/pleasure) and "systemization" (applying consistent principles over time) to ensure future generations do not view past actions as monstrous.
- Moral Parliament: To navigate ethical uncertainty, he employs a "moral parliament" approach, imagining multiple competing moral frameworks (e.g., utilitarianism, deontology, integrity-based) negotiating within the decision-maker to reach a compromise that avoids catastrophic outcomes for any single view.
- Rejection of "Ends Justify Means": Despite high stakes, Karnofsky rejects unethical tactics (lying, law-breaking, coercion), arguing that the integrity of the process is a non-negotiable constraint even in high-stakes scenarios.
- Stakeholder Management: He advocates for Open Philanthropy to remain small to maintain agility, noting that as organizations grow, they must accommodate more stakeholders, often reducing their ability to make disruptive or unconventional moves.
Forecasting and Methodology
- Karnofsky utilizes multiple data points for AI timeline predictions:
- Historical Analysis: Reviewing past long-term predictions, finding them unreliable but not hopeless.
- Biological Anchors: Comparing AI computational power to human brain capacity, estimating that AI capable of human-level reasoning will likely be affordable and feasible within this century.
- Expert Surveys: Acknowledging that AI researchers generally predict transformative capabilities within a few decades.
- Effort Investment: Noting that the majority of all AI effort in human history will likely occur in this century due to the massive recent scaling of resources.
- He rejects the idea that past prediction failures (e.g., Malthusian limits) prove future prediction is impossible, citing societal progress in making rigorous, data-driven forecasts.
- Karnofsky views the "innovation is mining" metaphor as inapplicable to AI, noting that AI success by one entity does not prevent success by others, unlike the discovery of natural resources.
- Karnofsky utilizes multiple data points for AI timeline predictions:
Comparisons and Disagreements
- Will McCaskill: Karnofsky finds McCaskill's book on long-termism too broad; he is "picky," focusing only on a short list of issues (primarily transformative AI) where he believes specific, predictable interventions are possible.
- Progress Studies: While sympathetic to the benefits of technological growth, he argues that humanity has reached a "grey zone" where the risks of unchecked growth (nuclear war, bioweapons, misaligned AI) now outweigh the historical benefits of pure acceleration.
- Open Philanthropy's AI Grant: He defends a $30 million grant to OpenAI as a strategic move to influence governance and secure a board seat, arguing it was a net positive despite the acceleration of AI capabilities.
Organizational Strategy and Communication
- Karnofsky maintains the "Cold Takes" blog as a public forum to stress-test his views, invite criticism, and attract talent to neglected areas like AI alignment and long-term forecasting.
- He admits that while the blog has led to deepened understanding and identification of weak points in his arguments, it has not yet caused a fundamental shift away from the core "Most Important Century" thesis.
- Open Philanthropy's strategy focuses on "seeding" new fields (like AI alignment) when traditional institutions are unwilling to fund them, aiming to build a permanent workforce of experts ready for the transition.