newsfilter.io
Interview

Sholto Douglas & Trenton Bricken — How LLMs actually think

  • Context and Reasoning Capabilities: Models with expanded context windows are projected to enable "meta-learning," allowing systems to master new languages in months and perform reasoning tasks with higher sample efficiency than humans, potentially acting superhumanly in specific information ingestion contexts.
  • Agent Reliability and Deployment: Widespread adoption of AI agents depends on resolving infrastructure speed issues and achieving a critical "step function" in reliability (an extra "nine") to support task chaining over hours or days, which will determine the automatability of entire job families.
  • Economic and Compute Constraints: The "intelligence explosion" may be dampened by the quadratic or dominant linear costs of attention and training, where achieving the next leap (e.g., GPT-7) could require orders of magnitude more compute, potentially limiting progress if models remain economically "stuck" without algorithmic breakthroughs.
  • Research Acceleration and Automation: AI is expected to significantly speed up the next few years of research by automating software engineering tasks, serving as a "fantastic co-pilot," and generating synthetic data with verified reasoning traces, though current teams remain "compute bound."
  • Neural Scaling and Feature Learning: The "Quantum Theory of Neural Scaling" suggests models learn features in a universal order (n-grams, then induction heads) and may exhibit "superposition," where features are compressed into high-dimensional spaces, making interpretability difficult without techniques like dictionary learning.
  • Interpretability and Safety: Future interpretability efforts may shift toward "automated interpretability" using models to debate features and find deception circuits, while synthetic data generation and "constraining rules" are viewed as crucial for safety and verifying model alignment.
  • Biological Parallels: AI research draws parallels to human cognition, positing that intelligence is based on a hierarchy of associative memories, that the cerebellum supports cognitive tasks via evolutionarily conserved attention mechanisms, and that models' "reconstructive memory" is analogous to human memory fabrication.
  • Capabilities and Limitations: Current models may struggle to mix concepts at human reasoning levels and are limited to a finite number of forward passes (approx. 5-7 recursion levels), though they may eventually achieve "superintelligence" by combining high reliability and long context rather than fundamental architectural changes.