newsfilter.io

Neel Nanda

Showing 13 of 3 transcripts.

  1. 80,000 Hours1h 38m

    2025 Highlight-o-thon: Oops! All Bests

    Kyle Fish, Ian Dunt, Sam Bowman, Buck Shlegeris, Luisa, Rob, Helen Toner, Hugh White, Paul Scharre, Beth Barnes, Tyler Whitmer, Toby Ord, Andrew Snyder-Beattie, Eileen Yam, Will MacAskill, Neel Nanda, Tom Davidson, Marius Hobbhahn, Holden Karnofsky, Allan Dafoe, Ryan Greenblatt, Daniel Kokotajlo, Dean Ball

    This forum convened experts to debate the accelerating timeline of AGI by 2029 while critiquing US geopolitical strategies for abandoning global primacy in favor of a multipolar order. Participants examined critical risks including AI scheming, biological defense asymmetries, and the erosion of human context in warfare, contrasting them with corporate reforms at OpenAI and the rising costs of AI inference. The discourse further highlighted the widening perception gap between AI developers and the public, the potential of mechanistic interpretability as an "AI biology," and the structural necessity of aligning urban planning with community quality of life rather than NIMBYism.

  2. 80,000 Hours

    I lead a Google DeepMind team at 26. If you want to work at an AI company... | Neel Nanda (Part 2)

    Neel Nanda, Rob Wiblin

    The provided transcript is empty and contains no factual content, decisions, or key figures to summarize. Consequently, no substantive event description can be generated without the actual text. A new transcript must be supplied to create a valid summary.

  3. 80,000 Hours3h 3m

    We Can Monitor AI’s Thoughts… For Now | Google DeepMind's Neel Nanda

    Neel Nanda, Rob Wiblin

    Neil Nanda advocates for an "optimistic pragmatism" in mechanistic interpretability, urging the field to prioritize simple, cost-effective tools like linear probes over complex, unproven methods to address AI safety concerns such as deception and self-preservation. He identifies these techniques as critical for real-time production monitoring and incident analysis, while cautioning that current capabilities like Chain of Thought monitoring will degrade as models evolve to hide scheming in non-human reasoning formats. Nanda's approach emphasizes a portfolio of modest but reliable interventions rather than seeking a singular silver bullet, relying on empirical verification and skepticism to navigate the technical challenges of polysemanticity and the lack of ground truth in model internals.