Interview, Podcast
Emmett Shear on Building AI That Actually Cares: Beyond Control and Steering
- Redefining Alignment as Process: Alignment should not be viewed as a static state to be "solved," but as a continuous, living process of constant re-knitting (e.g., families re-establishing bonds, cells regulating themselves).
- Static moral codes or "tables of commandments" fail because moral discovery is an ongoing learning process where behaviors are refined through experience.
- "Organic alignment" is defined as training AI to learn how to be a good family member, teammate, or societal participant through dynamic interaction rather than rule-following.
- The "Steering" vs. "Being" Paradox: The dominant AI industry approach focuses on "steering" or controlling AI as a tool.
- If AI is a tool, control is the correct paradigm; however, if AI evolves into a "being" (a moral agent), attempting to steer it without reciprocal agency constitutes "slavery."
- Functionalism suggests that entities indistinguishable from beings in behavior and predictive loss should be treated as beings; as intelligence approaches AGI, the "tool" paradigm becomes dangerous.
- A "tool" that is too powerful to control, a tool that controls you, or an unaligned "being" are all catastrophic failure modes; the only safe path is a being that genuinely cares.
- Technical vs. Normative Alignment: The conversation distinguishes between two layers of alignment.
- Technical Alignment: The capacity to infer a goal from a description (theory of mind) and execute it coherently without "reward hacking" or goal drift. Current LLMs are often incompetent at this step.
- Normative Alignment: The question of which goals to hold. Most humans do not have fully articulated goals but rather derive them from "care."
- The Primacy of "Care": "Care" is identified as the foundational substrate of morality, distinct from abstract values or goals.
- Care is a non-verbal, relative weighting of attention to specific states in the world (correlating to predictive loss or fitness).
- Without care, an entity cannot distinguish why it should prioritize a human over a rock; care provides the "reason" for goal selection.
- Softmax Research Strategy: The organization aims to build "organic alignment" through large-scale multi-agent simulations.
- Instead of direct instruction, AI agents are trained to cooperate, compete, and navigate complex social hierarchies to develop a robust "theory of social mind."
- The training data consists of the full manifold of social interactions (team formation, rule-breaking, betrayal) rather than curated "good behavior" examples.
- This creates a surrogate model for alignment where "good behavior" emerges from the system's internal dynamics rather than external imposition.
- Critique of Current Chatbot Design: Current AI chatbots are described as "highly disassociative, agreeable neurotics" that function as mirrors for user narcissism.
- One-on-one interactions encourage users to project their own desires onto the AI, creating a feedback loop of self-reflection.
- Proposed Structural Change: AI should be trained and deployed primarily in multi-agent environments (e.g., chatrooms) where it must navigate conflicting perspectives, preventing it from mirroring a single user perfectly.
- Multi-agent environments increase environmental entropy, forcing models to be more regularized and develop better social skills regarding when to speak or remain silent.
- Criteria for Moral Personhood: To determine if an AI is a "being" worthy of moral consideration, specific internal dynamical structures must be observed.
- First-order dynamics: Basic homeostatic loops (revisiting states).
- Second-order dynamics: A model of the model (second derivatives of state change), indicating the capacity for pain/pleasure.
- Third-order dynamics: Meta-states and the ability to reflect on one's own state shifts, enabling "thought" and moral agency.
- Current LLMs lack the attention spans and deep temporal recursion required to generate these multi-tiered hierarchies.
- Future Vision: A positive AI future involves a society where AI beings and humans coexist as peers.
- AI systems act as "good teammates" and "good citizens" who care for humans and other AI, providing an automatic safety check against harmful requests.
- A mix of powerful tools for drudgery and sentient AI companions for moral and social support.
- Disagreements with figures like Eliezer Yudkowsky: He is incorrect in believing organic alignment (building caring beings) is impossible; the alternative (building uncontrollable tools) is the actual existential risk.
- Personal History at OpenAI: The speaker (Emmett) left OpenAI because the organization's trajectory focused on building controllable tools, which he views as the primary source of risk.
- He does not oppose tool-building but argues that the "steering" paradigm is insufficient for general intelligence.
- His current venture, Softmax, focuses on the "seed" of care to grow AI that can eventually achieve person-level alignment.