Interview, Fireside Chat
AI progress is about to rapidly accelerate in 2025 – Sholto Douglas & Trenton Bricken
- Current Research Bottlenecks
- Progress is currently constrained more by compute availability and the "taste" required for inference on imperfect information than by pure engineering effort.
- "Taste" refers to the ability to make difficult strategic decisions and identify effective directions without complete data.
- Compute Elasticity and Scaling
- The Gemini program demonstrates an elasticity of approximately 0.5, estimating that a 10x increase in compute (e.g., H100s) would yield roughly a 5x increase in research speed and effectiveness.
- Increased compute converts directly to progress, with the primary constraint being the physical limit of available hardware (e.g., Sam 7 trillion parameters) rather than current intelligence levels.
- Teams must strategically allocate fixed compute between:
- Inference and experimentation: Running new research trials to test hypotheses.
- Scaling runs: Continuing to train the latest best models to capture emergent properties not visible in smaller scales.
- Excessive focus on research efficiency without continued investment in frontier-scale training risks derailing the long-term trajectory of the technology.
- Nature of the Intelligence Explosion
- An intelligence explosion is not viewed as AI writing code from scratch but rather as AI augmenting top researchers to significantly accelerate algorithmic progress.
- Immediate Mechanism: AI acts as a "fantastic co-pilot" for coding, enabling faster iteration on sub-tasks and sub-goals.
- Future Mechanism: Progress may eventually rely on "synthetic data" generated by AI as a crucial ingredient for model capability.
- AI may reach a point where it is faster for researchers to onboard and direct AI agents than to perform tasks personally.
- The Research Cycle and Decision Making
- The core research workflow involves:
- Generating a long list of potential ideas.
- Proving out concepts at various scales.
- Interpreting failures and understanding underlying mechanisms.
- A significant portion of research involves "introspection" to diagnose why specific ideas failed, rather than merely executing code.
- Imperfect Information:
- Researchers must predict whether performance trends observed at smaller scales will hold for larger architectures; this is not guaranteed.
- Features that improve performance at small scales can sometimes hinder performance at larger scales.
- Technical reports often omit the "graveyard" of failed intermediate runs, showing only smooth success curves.
- The core research workflow involves:
- Key Success Factors for Researchers
- Ruthless Prioritization: Distinguishing quality research from stagnation requires aggressively filtering which ideas to pursue under uncertainty.
- Simplicity Bias: Effective researchers avoid attachment to familiar academic toolboxes, instead attacking problems directly with expanded methodologies.
- Rapid Iteration: The ability to quickly cycle through experimentation, results, interpretation, and communication is the primary differentiator between high-performing teams.
- Cross-Disciplinary Tools: Top engineers integrate concepts from reinforcement learning, optimization theory, and systems engineering.
- Implications for Model Development
- The empirical nature of machine learning research is driving solutions toward "brain-like" architectures.
- The field is effectively performing "greedy evolutionary optimization" over the landscape of possible architectures and designs.