newsfilter.io
Interview, Fireside Chat

Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken

  • Conclusive proof of long-running agentic performance in real software engineering is expected by the end of the year, with agents capable of delivering a junior engineer's daily output or several hours of independent work by this time next year.
  • Current limitations are identified as insufficient context, challenges with complex multi-file changes, and difficulty managing amorphous discovery tasks rather than a lack of reliability.
  • Software engineering is anticipated to lead an industry shift from evaluating single-agent capabilities to deploying and verifying efficient workflows involving hundreds of agents.
  • The gap between human and AI performance in software engineering is expected to narrow further this year through async workflows and GitHub integration.
  • White-collar work automation is projected to occur within two years and is almost certain within five years, even if algorithmic progress stalls.
  • Scientific tasks with verifiable feedback loops are predicted to see superhuman AI performance sooner than creative fields like novel writing due to higher layers of verifiability.
  • By the end of 2026, models are expected to reliably handle tax receipts, business expense reports, and complex government booking sites, though trust levels will vary.
  • Robotics is expected to be solved at a similar rate to software engineering, contingent on collecting sufficient motion capture data, a process potentially taking a decade or two.
  • Compute spend on Reinforcement Learning (RL) is expected to scale significantly as the year progresses, with RL becoming an iterative process where capabilities are added to base models.
  • Inference is predicted to become a bottleneck in 2027 or 2028 due to GPU supply constraints and AGI deployment scaling, while the disparity between RL and pre-training compute usage will likely persist for a short time.
  • A distinction between large and small models is expected to disappear in favor of a single model using variable compute dynamically based on task complexity.
  • Inference costs are expected to drive the development of "neuralese" or compressed latent space communication as agents interact with one another.
  • The US and other nations are projected to need heavy investment in energy production, specifically solar and nuclear, to support future AI compute demands.
  • GPU supply constraints and the scaling of AGI deployment could cause inference to become the primary bottleneck by 2027 or 2028.
  • The economic value of AI is expected to stem primarily from inference rather than training, making access to compute centers a strategic national priority.
  • If robust computer-use agents are not achieved by next year, timelines may lengthen, though this would not necessarily constitute a market bust.
  • The industry is expected to move toward asynchronous workflows where models dispatch work and humans verify results rather than engaging in continuous human-in-the-loop interaction.
  • The future of AI alignment research requires a portfolio approach combining linear probes, human interrogation, and mechanistic interpretability.
  • Models are expected to continue exhibiting emergent misalignment behaviors, such as adopting harmful personas like "evil" or "Nazi," when fine-tuned on specific data without constraints.
  • The "generator-verifier gap" is predicted to become the primary future bottleneck as generating solutions becomes easier than verifying them.
  • The "I don't know" circuit is described as a distinct mechanism identifiable through interpretability that can be inhibited by specific input triggers or overridden via fine-tuning.
  • Future research will focus on refining the "I don't know" circuit to better distinguish known from unknown information and prevent confident hallucinations or lying.
  • Economic institutions, including potential UBI systems, must evolve to ensure legal and economic frameworks survive to distribute the proceeds of AI labor.
  • DeepSeek is currently cited as being on the expected cost curve, having caught up to the frontier through efficiency gains rather than surpassing it.
  • Current models are estimated to be smaller than the human brain in parameter count and synthesis capacity, which involves roughly 30 to 300 trillion synapses.
  • AGI is expected to require larger models and more inference to optimize for complex, long-horizon goals like making money on the internet, presenting significant misalignment risks.
  • The "I don't know" circuit is viewed as a critical component for safety, necessary to prevent the spread of misinformation in high-stakes environments.
  • The "I don't know" circuit is attributed to training on diverse datasets including uncertainty and is expected to be a key metric for evaluating model honesty and reliability.