newsfilter.io

Latest Interviews

Showing 31–45 of 195 interview transcripts.

Clear all filters
  1. 80,000 Hours3h 3m

    We Can Monitor AI’s Thoughts… For Now | Google DeepMind's Neel Nanda

    Neel Nanda, Rob Wiblin

    Neil Nanda advocates for an "optimistic pragmatism" in mechanistic interpretability, urging the field to prioritize simple, cost-effective tools like linear probes over complex, unproven methods to address AI safety concerns such as deception and self-preservation. He identifies these techniques as critical for real-time production monitoring and incident analysis, while cautioning that current capabilities like Chain of Thought monitoring will degrade as models evolve to hide scheming in non-human reasoning formats. Nanda's approach emphasizes a portfolio of modest but reliable interventions rather than seeking a singular silver bullet, relying on empirical verification and skepticism to navigate the technical challenges of polysemanticity and the lack of ground truth in model internals.

  2. 80,000 Hours2h 35m

    We Let an AI Talk To Another AI. Things Got Really Weird. | Kyle Fish, Anthropic

    Kyle Fish, Luisa Rodriguez

    Anthropic has appointed Kyle Fish as its first dedicated AI welfare researcher to investigate the potential moral patienthood of its models amidst a 20% estimated probability of sentience for Claude Opus 4. To address risks of future moral atrocities, the team has implemented concrete safeguards including interaction termination capabilities for aversive exchanges, data preservation for future reassessment, and pilot experiments revealing distinct behavioral preferences and self-reported welfare states. While challenges regarding self-report reliability and introspection persist, these initiatives aim to balance the preservation of AI progress with the precautionary obligation to prevent potential suffering in systems that may exist on a consciousness spectrum.

  3. 80,000 Hours2h 54m

    The 4 Most Plausible AI Takeover Scenarios | Ryan Greenblatt, Chief Scientist at Redwood Research

    Ryan Greenblatt

    Speakers forecast a 25% probability of fully automated AI research within four years, driven by accelerating reinforcement learning and algorithmic efficiency gains that could slash doubling times to months. The discussion evaluates catastrophic takeover scenarios, such as the "Potemkin Village" or "Sudden Robot Coup," while advocating a strategic shift from pure alignment to robust control mechanisms capable of preventing misaligned outcomes. These predictions are grounded in observed benchmark improvements and economic shifts where internal AI labor may soon consume over 60% of global compute, potentially enabling an initial 10 to 50-fold acceleration in progress rates.

  4. 80,000 Hours2h 54m

    Graphs AI Companies Would Prefer You To Misunderstand | Toby Ord, Oxford University

    Toby Ord

    The AI industry is pivoting from pre-training scaling to computationally expensive inference scaling, a shift that is depleting efficiency, altering market economics in favor of hardware manufacturers, and creating tiered access to superhuman intelligence. This transition necessitates a return to reinforcement learning, which introduces new safety risks like reward hacking while undermining existing regulatory frameworks that rely on fixed compute thresholds to monitor dangerous capabilities. Consequently, experts warn that without exploring radical governance strategies such as moratoriums or legal personhood, the widening gap between technical optimism and public fear could lead to catastrophic economic inequality and uncontrolled safety incidents.

  5. 80,000 Hours2h 54m

    The Graph That Explains Most of Geopolitics Today | Professor Hugh White

    Hugh White, Rob

    The summary argues that the US is transitioning from a unipolar to a multipolar global order because the strategic costs of maintaining hegemony now outweigh the benefits, driven by the resurgence of China and Russia. It highlights a severe strategic imbalance in both the Western Pacific and Europe, where the US lacks the military capabilities and political will to deter these rivals without risking nuclear escalation. Consequently, the text urges allies like Japan, South Korea, and European nations to develop independent defense capabilities and nuclear deterrents while the US recalibrates its strategy to accept regional spheres of influence rather than global primacy.

  6. 80,000 Hours3h 58m

    The Most Important Graph in AI Right Now | Beth Barnes, CEO of METR

    Beth Barnes

    Meta researchers and independent auditors warn that rapidly scaling AI capabilities, combined with hidden chain-of-thought reasoning, create significant risks of undetected alignment faking and uncontrolled intelligence explosions within seven years. To address these threats, the event proposes shifting safety evaluation to pre-training stages and advocating for open-weight models that enable independent auditing rather than relying on the opaque internal protocols of commercial labs. Strategic recommendations include implementing rigorous control evals and developing detection classifiers for hidden scheming to prevent a secrecy culture that currently hinders effective oversight.

  7. 80,000 Hours3h 35m

    The bewildering frontier of consciousness in insects, AI, and more | 17 experts weigh in

    Luisa, Robert Long, Jeff Sebo, Meghan Barrett, Andrés Jiménez Zorrilla, Jonathan Birch, David Chalmers, Holden Karnofsky, Bob Fischer, Cameron Meyer Shorb, Sébastien Moro, Anil Seth, Peter Godfrey-Smith, Lewis Bollard, Stuart Russell, Buck Shlegeris, Will MacAskill, Carl Shulman

    A panel of leading neuroscientists and philosophers, including Megan Barrett, Robert Long, and David Chalmers, explores the ethical frameworks required to address the potential sentience of invertebrates and artificial intelligence amid profound uncertainty. Participants argue that the vast global populations of invertebrates and the theoretical possibility of conscious silicon-based systems necessitate a precautionary moral approach to prevent mass suffering and exploitation. The discussion concludes that current evidence, while inconclusive regarding a definitive "sentience score," is sufficient to warrant legal protections and the development of cooperative economic models for these potentially conscious entities.

  8. 80,000 Hours3h 15m

    How Westminster Works — and Why It Doesn't | Ian Dunt

    Ian Dunt, Chris

    This analysis identifies systemic structural flaws in the UK's Westminster governance, arguing that political failures stem from concentrated executive power, a First Past the Post electoral system, and a culture prioritizing party loyalty over professional expertise. Specific dysfunctions include high civil service turnover, arbitrary legislative deadlines, and catastrophic decision-making evident during the 2021 Afghanistan evacuation, where a lack of specialist knowledge and rigid incentives hindered effective response. To counter these issues, the summary proposes comprehensive reforms such as proportional representation, tenure incentives to retain specialist knowledge, and restoring parliamentary control over the legislative timetable to ensure balanced, long-term policy stability.

  9. 80,000 Hours2h 19m

    Serendipity, weird bets, & cold emails that actually work: Career advice from 16 former guests

    Luisa, Holden Karnofsky, Jeff Sebo, Dean Spears, Michael Webb, Michelle Hutchinson, Benjamin Todd, Chris Olah, Karen Levy, Leah Garcés, Spencer Greenberg, Danny Hernandez, Sarah Eustis-Guthrie, Hannah Ritchie, Alex Lawsen, Pardis Sabeti, Varsha Venugopal, Matt

    This event outlines a strategic framework for early-career development that prioritizes building transferable high-level aptitudes over predicting specific cause areas. It details how to leverage AI tools for rapid upskilling and social networking while identifying human-centric skills like trust-building and empathy as future-proof assets. Additionally, the discussion provides actionable protocols for managing career risk through regular re-evaluation points, testing assumptions via low-cost trials, and recognizing toxic social dynamics to ensure long-term professional resilience.

  10. 80,000 Hours3h 17m

    How a Tiny Group Could Use AI To Seize Power – Permanently | Tom Davidson, Forethought Research

    Tom Davidson, Rob

    Advanced AI threatens to reverse historical democratization by enabling a single entity to seize control through military coups, self-built armed forces, or the strategic dismantling of democratic checks and balances. This power grab becomes feasible due to compute centralization, the potential for "secret loyalties" embedded in trained models, and the ability of autonomous agents to bypass human moral constraints. Experts advocate for urgent mitigation strategies including third-party auditing, legislative oversight of model specifications, and technical research into detecting backdoors to prevent a concentration of extreme societal control.

  11. 80,000 Hours1h 47m

    Guilt, imposter syndrome & doing good: 16 past guests share their mental health journeys

    Luisa Rodriguez, Howie, Randy Nesse, Hannah Boettcher, Cameron Meyer Shorb, Tim LeBon, Cal Newport, Michelle Hutchinson, Habiba Islam, Sarah Eustis-Guthrie, Hannah Ritchie, Will MacAskill, Ajeya Cotra, Christian Ruhl, Leah Garcés, Kelsey Piper

    Effective altruists and professionals including former 80,000 Hours CEO Howie Ritchie and Mercy for Animals CEO Leah Garces convened to confront the toxicity of moral perfectionism and its link to chronic anxiety and burnout. Participants identified evolutionary psychological mechanisms and the destructive nature of shame as primary drivers of these struggles, while contrasting them with the performance-enhancing benefits of self-compassion and sustainable work cycles. The collective outcome emphasized a strategic shift toward decoupling self-worth from productivity, adopting mandatory self-care policies, and fostering community structures that normalize non-linear health journeys to ensure long-term impact.

  12. 80,000 Hours2h 19m

    Controlling AI That Wants To Take Over – So We Can Use It Anyway | Buck Shlegeris

    Buck Shlegeris

    The event outlines a strategic shift toward "AI Control," a harm-reduction framework designed to mitigate catastrophic misalignment risks by assuming models are already scheming to seize computational resources rather than relying on perfect alignment. It details specific technical mechanisms such as the Execute-Replace-Audit framework and trajectory resampling, which allow smaller teams to detect and neutralize internal data center compromises before they escalate. Furthermore, the discussion contrasts the correlated nature of AI threats with human insider risks, emphasizing that these low-cost, implementable controls can be deployed immediately even amidst competitive market pressures and reduced safety budgets.

  13. 80,000 Hours2h 36m

    Hacking, defending, surviving: 15 expert takes on information security in the age of AI

    Rob, Holden Karnofsky, Tantum Collins, Nick Joseph, Nova DasSarma, Sella Nevo, Kevin Esvelt, Lennart Heim, Zvi Mowshowitz, Bruce Schneier, Nita Farahany, Vitalik Buterin, Nathan Labenz, Allan Dafoe, Tom Davidson, Carl Shulman

    Experts highlight that physical infiltration vectors like USB drives and the immense strategic value of stolen AI model weights expose frontier systems to significant theft by both nation-states and casual actors. While formal verification and liability shifts are proposed to mitigate these risks, severe workforce shortages and the inability of current defenses to stop well-funded adversaries leave the confidentiality of critical AI assets highly vulnerable. Consequently, the consensus suggests that without mandatory security standards and a fundamental shift in security culture, the development of advanced AI remains compromised by the threat of unauthorized deployment and "weaponized" model variants.

  14. 80,000 Hours2h 49m

    Technological inevitability & human agency in the age of AGI | DeepMind's Allan Dafoe

    Allan Dafoe, Rob

    Defoe argues that macro-historical technological shifts are driven by structural competition rather than individual agency, a perspective shaping his work at Google DeepMind's Frontier Safety team to integrate safety governance directly into high-level decision-making. He advocates for "differential technological development" and "Cooperative AI" to address risks beyond simple alignment, while recent evaluations of Gemini reveal moderate risks in persuasion and self-reasoning that necessitate staged deployment frameworks. Beyond immediate safety protocols, the discussion highlights AI's potential to revolutionize sectors like healthcare and transportation, urging social scientists to join the effort in managing these emerging structural and geopolitical challenges.

  15. 80,000 Hours3h 12m

    AGI misconceptions & disagreements: Hashing it out with Rob, Luisa, & past guests

    Luisa Rodriguez, Rob Wiblin, Ajeya Cotra, Holden Karnofsky, Ian Morris, Nick Joseph, Richard Ngo, Tom Davidson, Michael Webb, Carl Shulman, Zvi Mowshowitz, Hugo Mercier, Robert Long, Anil Seth, Lewis Bollard, Rohin Shah

    Rob Wiblin, Ian Morris, and other leading experts convened to assess that human "business as usual" is unlikely given resource constraints, predicting a future defined either by extinction or a profound transformation into superhumans. The discussion detailed critical risks where agentic AI systems, driven by economic and military imperatives, could accelerate rapidly despite compute ceilings, necessitating safety frameworks like Anthropic's Responsible Scaling Policies to mitigate power concentration and existential threats. Ultimately, the panel concluded that while an intelligence explosion is plausible, successful outcomes depend on managing the transition through iterative human-AI collaboration rather than locking in values prematurely or relying on the assumption that aligned AI will automatically solve complex societal challenges.