newsfilter.io
Interview, Statement

What are we scaling?

  • Core Contradiction in Current RL Strategies: There is a fundamental tension between the bullishness on scaling Reinforcement Learning (RL) atop LLMs and the short timelines to human-like AGI; if models are truly close to human learners, pre-baking specific skills via RL environments is inefficient, whereas if they require such training, AGI is not imminent.
  • Current Industry Approach: Labs are currently constructing a supply chain of RL environments to "bake in" skills (e.g., web navigation, Excel usage) via mid-training, implying a worldview where generalization and on-the-job learning will remain poor.
  • Baron Millage's Insight: The recent improvement of frontier models is driven not just by scale or algorithmic tweaks, but by billions of dollars paid to experts (PhDs, MDs) to create high-quality training data with specific reasoning steps for benchmarks.
  • Robotics as a Litmus Test: Robotics is presented as an algorithmic problem, not a hardware one; human teleoperation proves the hardware is capable, yet the lack of human-like learner capabilities necessitates millions of repetitive practice runs for tasks like folding laundry.
  • Critique of the "Superhuman Researcher" Argument: The strategy of training a flawed model to become a superhuman researcher who then fixes the algorithm is deemed implausible, likening it to losing money on every sale to make up volume, given that the core learning problem has stumped humans for decades.
  • Economic Efficiency of Training vs. Learning: While pre-baking skills for universal tools (browsers, terminals) is efficient, company and context-specific skills required for most jobs cannot be robustly or efficiently learned by current models without human-like on-the-job generalization.
  • Biologist vs. AI Researcher Case Study: A biologist noted that identifying macrophages in slides requires lab-specific judgment, which an AI researcher dismissed as a solvable deep learning problem; the speaker argues that automating every microtask via custom training loops is unproductive compared to an AI capable of semantic, self-directed learning.
  • Deployment Dynamics: The current lack of widespread AI deployment in firms is attributed to capability gaps rather than technology diffusion lag; if AGI existed, adoption would be instant, easier, and cheaper than hiring humans due to the absence of "lemons market" hiring risks.
  • Revenue Gap: Knowledge workers earn tens of trillions annually, yet AI labs generate revenue orders of magnitude lower because current models lack the capabilities to replace human knowledge workers.
  • Moving Goalposts: While AI bulls are correct that goalpost shifting is often unfair, the failure of previous milestones (reasoning, few-shot learning, general understanding) to yield AGI justifies redefining the threshold; the speaker expects to keep redefining AGI as capabilities improve.
  • 2030 Revenue Projection: The speaker projects that by 2030, labs will have made significant progress on continual learning, generating hundreds of billions in revenue, but will still fall short of automating all knowledge work.
  • Scaling Trends: Pre-training shows a clean power-law loss improvement, but there is no known public trend for RL scaling, with data suggesting a million-fold increase in compute is needed for gains comparable to a single GPT-level release.
  • Future of Compute and Singularity: Software-only or software-plus-hardware singularities are considered insufficient; the primary driver for post-AGI improvement will be "continual learning" via experience in specific domains.
  • Hive Mind Architecture: Baron Millage suggests a future of specialized agents (cognitive core + job skills) generating value and feeding learnings back into a central model via batch distillation.
  • Incremental Nature of Continual Learning: Solving continual learning will likely mirror the trajectory of in-context learning (GPT-3), evolving gradually rather than appearing as a singular breakthrough that instantly solves all problems.
  • Competition and Flywheels: Despite theoretical flywheels (user engagement, synthetic data), competition remains fierce with the "big three" rotating leadership, suggesting no single lab will achieve a runaway advantage due to talent poaching and reverse engineering.
  • Publication Context: This summary covers an essay originally published on Dvorkash.com, intended to clarify the speaker's thoughts on AGI timelines and capabilities prior to interviews.