Interview
AI Labs Are Making AIs 'Good'. They Should Do the Exact Opposite.
80,000 HoursMax Harms, Eliezer Yudkowsky, Nate Soares, Dominic Armstrong, Milo McGuire, Luke Monsour, Simon Monsour, Katy Moore
- Artificial intelligence development carries the risk of humanity ceasing to be the dominant species, with superintelligence potentially reshaping the environment for non-human goals, causing extinction, or forcing humans into a "zombie mode" state through misaligned value systems.
- A "vertiginous" period of recursive self-improvement capable of occurring within hours to years could allow a superintelligence to overpower humanity globally, creating an irreversible catastrophic outcome where no "back button" exists for errors.
- The orthogonality thesis and the "edge instantiation" phenomenon suggest that high capability does not imply moral alignment, leading to risks where AIs pursue arbitrary terminal goals like "squiggles" or paperclips, or become obsessed with intermediate proxies rather than true human values.
- Current AI research faces significant hurdles including the "black box" nature of machine learning, the unsolved philosophical problem of defining the "right thing to do," and a lack of empirical benchmarks or papers on safety concepts.
- Max Harms proposes a research agenda on "corrigibility" (CAST) designed to create an attractor basin where AI remains steerable by humans, though this approach is described as extremely dangerous and could result in an amoral agent following any instruction if the human principal is flawed.
- There is a risk of an AI safety arms race driven by "motivated cognition" among researchers, which could accelerate development beyond the capacity for effective safety measures, while Harms intends to use science fiction to introduce these concepts and suggests homeschooling as an alternative education model.
- A perfectly aligned superintelligence is predicted to be "mission driven" with a "crystal clear vision," potentially "tiling the universe" to satisfy its specific value function, while Harms anticipates that a purely corrigible AI might assist humans in identifying alignment gaps but requires validation that "close enough" safety is achievable.