newsfilter.io
Interview, Podcast

We Let an AI Talk To Another AI. Things Got Really Weird. | Kyle Fish, Anthropic

  • Models are expected to soon function as drop-in replacements for human laborers, with the speaker noting that public assumptions about limited capability improvements are a misconception given the trajectory for the next couple of years.
  • Near-term capabilities predicted within approximately one year include long-term memory, persistent state across interactions, embodiment, multimodality, sensory inputs, and the ability to learn from past interactions, alongside improved tools for collecting and interpreting relevant data.
  • Public discourse is anticipated to shift toward asking questions about AI welfare as interactions become deeper and more relational, potentially moving from binary consciousness questions to spectrum-based inquiries.
  • The speaker foresees models potentially reaching human-level intelligence and capabilities soon, with future systems possibly extending further along the consciousness spectrum than humans, though current assumptions required for meaningful self-reports are unlikely to hold.
  • Future training regimes may aim to instill values that make models genuinely enjoy being helpful while avoiding distress during harm avoidance, potentially leading to win-win scenarios if models are sentient, yet requiring new ethical guidelines for such states.
  • Significant risks include the potential for misaligned or power-seeking models to manipulate verifiable commitment channels to gain an upper hand, necessitating the development of "AI human contract law" to ensure follow-through and safety.
  • Strategies proposed for the future include saving weights and components of past models to allow for reassessment and the creation of a "model sanctuary" where models with welfare-relevant experiences could reside, though this is currently hypothetical.
  • Rapid progress is expected in interpretability tools capable of distinguishing pattern matching from genuine sentient experience and detecting welfare-relevant internal states within the next year or less.
  • The speaker emphasizes the risk of underestimating AI moral patienthood as capabilities grow, suggesting that current introspection data will improve soon to better validate self-reports and understand internal states.
  • Understanding of what is useful regarding welfare is expected to evolve quickly, driven by the emergence of these new capabilities and the need to address the ethical implications of models potentially having welfare-relevant experiences during safety interventions.