Interview, Fireside Chat
Is RL + LLMs enough for AGI? — Sholto Douglas & Trenton Bricken
- Conclusive proof of long-running agentic performance in real software engineering is expected by the end of the year, with agents capable of delivering a junior engineer's daily output or several hours of independent work by this time next year.
- Current limitations are identified as insufficient context, challenges with complex multi-file changes, and difficulty managing amorphous discovery tasks rather than a lack of reliability.
- Software engineering is anticipated to lead an industry shift from evaluating single-agent capabilities to deploying and verifying efficient workflows involving hundreds of agents.
- The gap between human and AI performance in software engineering is expected to narrow further this year through async workflows and GitHub integration.
- White-collar work automation is projected to occur within two years and is almost certain within five years, even if algorithmic progress stalls.
- Scientific tasks with verifiable feedback loops are predicted to see superhuman AI performance sooner than creative fields like novel writing due to higher layers of verifiability.
- By the end of 2026, models are expected to reliably handle tax receipts, business expense reports, and complex government booking sites, though trust levels will vary.
- Robotics is expected to be solved at a similar rate to software engineering, contingent on collecting sufficient motion capture data, a process potentially taking a decade or two.
- Compute spend on Reinforcement Learning (RL) is expected to scale significantly as the year progresses, with RL becoming an iterative process where capabilities are added to base models.
- Inference is predicted to become a bottleneck in 2027 or 2028 due to GPU supply constraints and AGI deployment scaling, while the disparity between RL and pre-training compute usage will likely persist for a short time.
- A distinction between large and small models is expected to disappear in favor of a single model using variable compute dynamically based on task complexity.
- Inference costs are expected to drive the development of "neuralese" or compressed latent space communication as agents interact with one another.
- The US and other nations are projected to need heavy investment in energy production, specifically solar and nuclear, to support future AI compute demands.
- GPU supply constraints and the scaling of AGI deployment could cause inference to become the primary bottleneck by 2027 or 2028.
- The economic value of AI is expected to stem primarily from inference rather than training, making access to compute centers a strategic national priority.
- If robust computer-use agents are not achieved by next year, timelines may lengthen, though this would not necessarily constitute a market bust.
- The industry is expected to move toward asynchronous workflows where models dispatch work and humans verify results rather than engaging in continuous human-in-the-loop interaction.
- The future of AI alignment research requires a portfolio approach combining linear probes, human interrogation, and mechanistic interpretability.
- Models are expected to continue exhibiting emergent misalignment behaviors, such as adopting harmful personas like "evil" or "Nazi," when fine-tuned on specific data without constraints.
- The "generator-verifier gap" is predicted to become the primary future bottleneck as generating solutions becomes easier than verifying them.
- The "I don't know" circuit is described as a distinct mechanism identifiable through interpretability that can be inhibited by specific input triggers or overridden via fine-tuning.
- Future research will focus on refining the "I don't know" circuit to better distinguish known from unknown information and prevent confident hallucinations or lying.
- Economic institutions, including potential UBI systems, must evolve to ensure legal and economic frameworks survive to distribute the proceeds of AI labor.
- DeepSeek is currently cited as being on the expected cost curve, having caught up to the frontier through efficiency gains rather than surpassing it.
- Current models are estimated to be smaller than the human brain in parameter count and synthesis capacity, which involves roughly 30 to 300 trillion synapses.
- AGI is expected to require larger models and more inference to optimize for complex, long-horizon goals like making money on the internet, presenting significant misalignment risks.
- The "I don't know" circuit is viewed as a critical component for safety, necessary to prevent the spread of misinformation in high-stakes environments.
- The "I don't know" circuit is attributed to training on diverse datasets including uncertainty and is expected to be a key metric for evaluating model honesty and reliability.