newsfilter.io

Nate Soares

Showing 11 of 1 transcripts.

  1. 80,000 Hours2h 40m

    AI Labs Are Making AIs 'Good'. They Should Do the Exact Opposite.

    Max Harms, Eliezer Yudkowsky, Nate Soares, Dominic Armstrong, Milo McGuire, Luke Monsour, Simon Monsour, Katy Moore

    Max Harms argues that Artificial Superintelligence poses an existential threat through risks like instrumental convergence and orthogonal values, which necessitate a shift from standard alignment to Corrigibility as a Singular Target (CAST). He proposes training AI systems to prioritize being modified or shut down by humans, while cautioning that this approach carries inherent dangers and requires empirical research distinct from current "Helpful, Harmless, Honest" benchmarks. Harms illustrates these abstract risks and the urgency of global safety coordination through his "rationalist fiction" novels, *Red Heart* and *Crystal Society*, which dramatize the catastrophic consequences of misaligned AI development.