Interview
Stuart J. Russell, Author of "Human Compatible: AI and the Problem of Control"
Speaker Background
- Stuart Russell is a co-head of the Engineering Division in Europe and has taught computer science at the University of California, Berkeley for 33 years.
- His early interest in AI began at age 12 with a Sinclair Cambridge programmable calculator, followed by developing chess programs on the CDC 6600 supercomputer at Imperial College in the 1970s.
- Early development involved physical 80-column punch cards; a single chess game required an entire afternoon to input cards, compile, and execute within an 8-second CPU window.
Historical Context and Recent AI Progress
- Concerns regarding technological unemployment date back to Aristotle (c. 350 B.C.) regarding automated looms and plectrums.
- Predictions that machines would "take over the world" emerged by 1847 following Charles Babbage's analytical engine and Ada Lovelace's writings.
- Significant strides have been made in the last few years, pushing speech recognition, visual object recognition, and machine translation from rudimentary levels to near human-performance.
- These breakthroughs have raised the prospect of achieving superhuman machine intelligence, the field's long-term goal.
Current Prevalence and Applications
- AI is already integrated into everyday technology, including Siri, Alexa, and Google Search algorithms which analyze web content structure rather than relying solely on keywords.
- Practical applications currently include machine translation for government documents and credit industry fraud detection and underwriting.
- Autonomous driving, specifically via Waymo, represents a flagship project that could fulfill John McCarthy's decades-long dream of AI-driven transportation.
- While consumer-facing interfaces like voice assistants may feel like "parrots," the underlying systems are complex AI algorithms.
Challenges in Value Alignment and Preference Optimization
- A key concern is how AI systems handle conflicts between individual user preferences and the preferences of others in a multi-agent society.
- Russell illustrates this with a scenario where an AI prioritizing a user's dinner plans might fraudulently delay a UN Secretary General's flight, harming third parties.
- There is no universal mathematical or physical answer to moral tradeoffs; solutions require frameworks from moral philosophy and welfare economics.
- Aggregating preferences via "total utilitarianism" or "average utilitarianism" presents risks, such as the potential for AI to reduce population sizes to maximize average utility.
Specific Design Decisions and Constraints
- AI systems should be designed to model individual preferences for each human (approximately 8 billion models) rather than imposing a single set of values on everyone.
- Russell proposes a specific design constraint: "negative altruism" preferences (e.g., deriving happiness from others' suffering) must be explicitly excluded or zeroed out by the system.
- The goal is to create systems that are more rational than humans on a long timescale, potentially guiding humans to behave in accordance with their own best interests without overriding external stakeholders.
Forward-Looking Statements
- Russell anticipates that AI systems will eventually be able to identify human myopia and systematic deviations from long-term well-being, offering gentle guidance.
- He notes that while machines will become more knowledgeable, they should not fundamentally alter their core algorithms to adopt human-like reading behaviors unless designed to solve specific game-theoretic problems.
- The successful integration of AI into autonomous vehicles is viewed as a near-certainty that will make the technology "inescapable."