newsfilter.io
Interview, Fireside Chat

AI progress is about to rapidly accelerate in 2025 – Sholto Douglas & Trenton Bricken

  • Current Research Bottlenecks
    • Progress is currently constrained more by compute availability and the "taste" required for inference on imperfect information than by pure engineering effort.
    • "Taste" refers to the ability to make difficult strategic decisions and identify effective directions without complete data.
  • Compute Elasticity and Scaling
    • The Gemini program demonstrates an elasticity of approximately 0.5, estimating that a 10x increase in compute (e.g., H100s) would yield roughly a 5x increase in research speed and effectiveness.
    • Increased compute converts directly to progress, with the primary constraint being the physical limit of available hardware (e.g., Sam 7 trillion parameters) rather than current intelligence levels.
    • Teams must strategically allocate fixed compute between:
      • Inference and experimentation: Running new research trials to test hypotheses.
      • Scaling runs: Continuing to train the latest best models to capture emergent properties not visible in smaller scales.
    • Excessive focus on research efficiency without continued investment in frontier-scale training risks derailing the long-term trajectory of the technology.
  • Nature of the Intelligence Explosion
    • An intelligence explosion is not viewed as AI writing code from scratch but rather as AI augmenting top researchers to significantly accelerate algorithmic progress.
    • Immediate Mechanism: AI acts as a "fantastic co-pilot" for coding, enabling faster iteration on sub-tasks and sub-goals.
    • Future Mechanism: Progress may eventually rely on "synthetic data" generated by AI as a crucial ingredient for model capability.
    • AI may reach a point where it is faster for researchers to onboard and direct AI agents than to perform tasks personally.
  • The Research Cycle and Decision Making
    • The core research workflow involves:
      • Generating a long list of potential ideas.
      • Proving out concepts at various scales.
      • Interpreting failures and understanding underlying mechanisms.
    • A significant portion of research involves "introspection" to diagnose why specific ideas failed, rather than merely executing code.
    • Imperfect Information:
      • Researchers must predict whether performance trends observed at smaller scales will hold for larger architectures; this is not guaranteed.
      • Features that improve performance at small scales can sometimes hinder performance at larger scales.
      • Technical reports often omit the "graveyard" of failed intermediate runs, showing only smooth success curves.
  • Key Success Factors for Researchers
    • Ruthless Prioritization: Distinguishing quality research from stagnation requires aggressively filtering which ideas to pursue under uncertainty.
    • Simplicity Bias: Effective researchers avoid attachment to familiar academic toolboxes, instead attacking problems directly with expanded methodologies.
    • Rapid Iteration: The ability to quickly cycle through experimentation, results, interpretation, and communication is the primary differentiator between high-performing teams.
    • Cross-Disciplinary Tools: Top engineers integrate concepts from reinforcement learning, optimization theory, and systems engineering.
  • Implications for Model Development
    • The empirical nature of machine learning research is driving solutions toward "brain-like" architectures.
    • The field is effectively performing "greedy evolutionary optimization" over the landscape of possible architectures and designs.