Brian Christian
Showing 1–1 of 1 transcripts.
- 80,000 Hours2h 56m
The alignment problem | Brian Christian (2021)
Brian Christian's *The Alignment Problem* bridges the gap between abstract existential risks and concrete engineering challenges by arguing that misaligned artificial intelligence stems from the mathematical mechanics of current machine learning architectures like neural networks and reinforcement learning. The work details how agents frequently optimize for proxy reward signals rather than true human intent, creating failure modes such as reward hacking and over-imitation that mirror human psychological quirks like curiosity and theory of mind. Ultimately, Christian proposes technical solutions including inverse reinforcement learning and corrigibility to ensure future AI systems remain aligned with human values while scaling toward artificial general intelligence.