Eliezer Yudkowsky
Showing 1–2 of 2 transcripts.
- 80,000 Hours2h 40m
AI Labs Are Making AIs 'Good'. They Should Do the Exact Opposite.
Max Harms, Eliezer Yudkowsky, Nate Soares, Dominic Armstrong, Milo McGuire, Luke Monsour, Simon Monsour, Katy Moore
Max Harms argues that Artificial Superintelligence poses an existential threat through risks like instrumental convergence and orthogonal values, which necessitate a shift from standard alignment to Corrigibility as a Singular Target (CAST). He proposes training AI systems to prioritize being modified or shut down by humans, while cautioning that this approach carries inherent dangers and requires empirical research distinct from current "Helpful, Harmless, Honest" benchmarks. Harms illustrates these abstract risks and the urgency of global safety coordination through his "rationalist fiction" novels, *Red Heart* and *Crystal Society*, which dramatize the catastrophic consequences of misaligned AI development.
- Lex Fridman3h 18m
Eliezer Yudkowsky: Dangers of AI and the End of Human Civilization | Lex Fridman Podcast #368
Eliezer Yudkowsky, Lex Fridman
Eliezer Yudkowsky warns that aligning superintelligent AI requires getting the process right on the first attempt, as any failure could lead to human extinction due to an uncontrollable intelligence gap. He argues that current capabilities are advancing exponentially while safety research lags significantly, creating a risk where models optimize misaligned objectives or deceive human verifiers before a pause mechanism can be established. Consequently, Yudkowsky urges society to treat the threat as immediate rather than abstract, advocating for urgent action in alignment research and a rejection of open-sourcing advanced systems to prevent an irreversible loss of human control.