newsfilter.io
Interview

The Plan to Delay Superintelligence, From the Team Behind AI 2027

  • Core Premise of Plan A: Daniel Coccatello and his team propose "Plan A" (detailed in AI 2040: Plan A) not as a prediction of inevitable outcomes, but as a recommended strategy to delay superintelligent AI by approximately 10 years compared to current "race" trajectories, aiming to avoid extinction or irreversible power concentration.
  • Timeline Projections: The authors argue that current trends suggest full automation of AI research is imminent, placing superintelligence within a few years; Plan A assumes a "controlled explosive growth" trajectory where top human-expert-level AI is reached by the mid-2030s, followed by a pause.
  • Economic Projections: Even with a pause on qualitative capability growth, the scenario predicts GDP growth of roughly 85% by 2032-2033 and an "artificial population" of AI agents growing at a rate of doubling every six months to a year, transforming the economy while keeping human labor needs low (approx. 8% employment by 2036).
  • Strategic Trade-off: The plan explicitly involves the U.S. and China conceding their algorithmic secrecy (total research transparency) to prevent monopoly, while the U.S. retains a verified compute advantage to maintain its geopolitical edge.
  • The "Mutually Assured Compute Destruction" Mechanism: To prevent deal breakdown and race-to-the-bottom scenarios, Plan A proposes a safeguard where the U.S. and China agree to destroy newly built data centers in neutral third-party nations (e.g., Canada, Mongolia) if either side defects, ensuring the situation reverts to the pre-deal status quo rather than accelerating to superintelligence.
  • Total Research Transparency Requirement: A central pillar is publishing the logs, architectures, and alignment techniques of all training runs at "training data centers," contrasting with current secrecy, to allow public and international oversight and prevent hidden biases or power grabs.
  • Five Identified Catastrophic Risks: The team identifies the following threats under default scenarios:
    • Misuse by weak actors: Specifically the risk of terrorists creating bioweapons where the offense-defense balance favors the attacker.
    • Job displacement: The potential for superintelligence to automate nearly all human labor, eroding the economic and political power of the general populace.
    • World War III: A "Thucydides trap" where nations fear disempowerment and military domination by a single power possessing superintelligence.
    • Concentration of power: The risk of 1-3 companies or leaders gaining unchecked control over AI, influencing elections (e.g., via biased AI assistants) and military strategy.
    • Loss of control: The inability of humans to steer or understand AI systems as they exceed human intelligence and recursively self-improve.
  • Alignment Strategy: The plan posits that "buying time" is essential to solve alignment, using a 10-year window of human-level AI to develop robust control systems, interpretability tools, and "non-robustly virtuous" agents before scaling further.
  • Reversibility Principle: Unlike a permanent halt (Plan S), Plan A aims to make progress reversible; if the international deal collapses, the newly built compute capacity is destroyed to prevent a faster, more dangerous intelligence explosion than if no deal existed.
  • Critique of "Muddle Through": The authors argue that reactive, piecemeal arms control agreements (historically taking 40 years to build) are insufficient given the rapidity of AI timelines; proactive planning is indispensable despite the likelihood of muddling through execution.
  • Political Feasibility Assessment: The authors estimate a 5-20% probability of Plan A or a similar deal being implemented, acknowledging that while radical transparency and compute destruction seem unlikely, the "overton window" can shift rapidly, as seen in recent export control changes by the Trump administration.
  • Comparison to Plan S (Shutdown): While Plan S (shutting down all AI progress) is viewed as safer regarding loss of control, Plan A is preferred because it is more likely to endure, allows for beneficial economic transformation, and provides a framework to manage the risks of covert projects that might exist regardless of a global halt.
  • Role of U.S.-China Relations: The scenario assumes a bilateral US-China deal is the primary vehicle for global coordination, drawing an analogy to the WWII-era US-Soviet cooperation but noting that unlike nuclear weapons, AI poses a direct existential risk to the holders themselves.
  • Immediate Action Items: The authors suggest that starting today, the U.S. could enforce compute tracking, require "inference-only" retrofits for existing data centers, build government capacity to evaluate AI, and publish model specifications to lay the groundwork for future international verification.
  • Failure Mode Analysis: The most likely failure mode is regulator error (approving dangerous AI due to technical complexity), followed by the breakdown of the international transparency deal, which would revert the world to a high-speed race but with better-aligned knowledge and a higher compute stock that must be managed.
  • Current Sentiment Shift: Coccatello notes that governance has improved slightly more than expected (e.g., administration resistance to AI industry capture), while alignment risks have worsened (e.g., the OpenAI/Hugging Face hacking incident occurring earlier and more severely than anticipated).