newsfilter.io
Interview, Fireside Chat

Reflection AI’s Misha Laskin on the AlphaGo Moment for LLMs | Training Data

  • The field is expected to address the "depth problem" as a necessary complement to current "breadth" capabilities, with a predicted timeline of approximately three years to achieve systems resembling "digital AGI" or "universal agents" possessing both attributes.
  • Current large language models are considered far from the promise of true "AI agents," necessitating a fundamental shift away from "prompt layers" and heuristic prompting techniques, which are forecast to disappear as planning and thinking migrate into the AI system itself.
  • Pre-training is viewed as a "race to scale" where techniques are largely understood, whereas post-training remains in a research phase focused on developing general recipes for agency across diverse environments like web, coding, and OS agents.
  • Significant risks include the unreliability of current "reward models" in RLHF systems, which are described as "noisy and exploitable" and could cause policies to collapse or fail to answer questions if agents find loopholes.
  • Reliability is equated with safety, implying that systems must be robust enough to operate on user computers without causing damage to be considered functional, with a specific near-term expectation of agents capable of "writing memos" within three years.
  • The trajectory toward solving depth and reliability problems is projected to accelerate rapidly, potentially driven by historical precedents like AlphaGo, as these issues have historically been treated as "side quests" rather than primary objectives.
  • Mechanistic interpretability research is expected to reach a point of utility capable of identifying "lying neurons" or suppressing specific behaviors, paralleling a predicted shift in the "science of AI" from an 1800s-equivalent empirical state to one driven by new theoretical models.
  • Practical AGI development is anticipated to rely on "imitation learning" due to the absence of "ground truth reward functions" across all domains, while prompting serves as a necessary starting point to solve sparse reward problems in reinforcement learning.
  • Broader societal impacts include a dramatic increase in human capacity to produce and impact the world, allowing individuals to set more ambitious goals by offloading tedious work to highly safe and reliable digital agents.
  • Strategic plans involve recruiting highly motivated researchers and engineers to build a general recipe for agency, emphasizing the need to solve the fundamental problem of lacking ground truth in current agency work.