Interview
Government and society after AGI | Carl Shulman (Part 2)
- Carl Shulman argues that the pace of technological, industrial, and economic change will intensify as AI automates the process of improving itself and developing other technologies.
- The critical window for safety alignment expands from months to years if AI systems gain the capability to automate research, creating a much higher cost for delays in alignment efforts.
- Shulman opposes voluntary AI pauses at the current stage, noting that political momentum and the urgency to address existential risks (like AI takeover or undermining nuclear deterrence) are significantly higher once AI capabilities are more advanced.
- A unilateral voluntary pause by concerned entities could disproportionately benefit adversarial actors (companies or nations less concerned with AI risk) by shifting relative influence and reducing the leader's slack to negotiate safety standards.
- The optimal strategy for safety is a binding international agreement or a coordinated race between large international blocks, rather than a free-for-all or a premature voluntary pause.
- Superhuman AI advisors could have revolutionized the response to the COVID-19 pandemic by providing trustworthy, unbiased forecasts that override local incentives to hide outbreaks or delay testing.
- In the Chinese context, AI forecasting of the economic and political costs of an uncontained pandemic could have compelled local officials to report data immediately, overcoming bureaucratic incentives to minimize their own blame.
- In the US, AI advisors could have pressured the CDC to lift testing bans and supported "Operation Warp Speed" by clarifying that faster vaccine deployment was the optimal strategy for public health and re-election prospects.
- AI systems could have counteracted anti-vaccine sentiment by demonstrating empirically to conservative groups that vaccine hesitancy reduced their voter base and harmed their interests.
- Shulman predicts that AI will drastically improve "soft" domains like social science, forecasting, and political philosophy, potentially converging on truth in areas currently dominated by subjective judgment or bias.
- There is a risk that individuals or regimes might use AI to "lock in" values and prevent future reflection, effectively creating epistemic bubbles where AI confirms dogmatic views rather than challenging them.
- Authoritarian regimes (e.g., North Korea, China, Iran) could use AI to radicalize populations and manipulate loyalty indexes, potentially causing leaders to become detached from reality due to feedback loops of propaganda.
- To prevent coups, future AI security forces must be programmatically designed to reject illegal orders and defend the constitutional order, requiring broad, pluralistic oversight of their core motivations.
- Trust in AI security forces requires "constitutional AI" principles where machines are trained to recognize and refuse orders that violate democratic norms, even if those orders come from legitimate human superiors.
- Current military reliance on human loyalty is being eroded by automation; future stability depends on machines being programmed to defend the legal order rather than following the whims of a single executive.
- AI forecasting can be trained using "held-out" data (e.g., training on data up to 2021 to predict 2022 outcomes) and reinforcement learning based on predictive accuracy, though macro-scale forecasting is limited by the small number of independent historical data points.
- The "interpretability" of AI models will allow society to verify that AI predictions are based on truth-tracking concepts rather than political bias or common human fallacies, similar to how scientific consensus transcends individual ideology.
- Robust epistemic rules (e.g., pre-registering hypotheses, avoiding p-hacking) can be codified into AI training to create systems that remain truthful even under intense pressure to deceive.
- Shulman suggests that adversarial testing, where a weaker AI attempts to expose lies in a stronger model while adhering to strict rules of reasoning, could help adversarial nations (like the US and China) verify the honesty of each other's AI.
- Key infrastructure for trust includes "watchdog" organizations that systematically audit corporate models for bias and dishonesty, creating market incentives for companies to produce verifiably honest AI.
- The most scalable near-term business model involves eliminating AI hallucinations and improving source verification, with specific applications in financial trading, political forecasting, and government compliance.
- Listeners can contribute by filling out the 80,000 Hours AI Census to help match skills in technical research, policy, and operations with organizations working on AI safety.
- Shulman maintains a median expectation that AI will lead to a prosperous future with improved public epistemology, while acknowledging the personal and existential risks of a potential disaster.