newsfilter.io
Interview, Fireside Chat

Forecasting & the drivers of AI progress | Danny Hernandez (2020)

Forecasting and Organizational Impact

  • Forecasting AI progress requires distinguishing between current capabilities and speculative future claims to prevent experts and the public from talking past one another.
  • "Ideological Turing tests"—explaining expert viewpoints well enough that the expert agrees with your summary—are a critical method for internalizing uncertainty and avoiding "straw man" errors in decision-making.
  • Organizations benefit from a "numeracy" culture where every employee, not just leadership, makes quantitative forecasts to reduce internal politics and increase the efficiency of resource allocation.
  • Calibration training (learning to predict outcomes with accurate probabilities) has been shown to increase meeting satisfaction from 50% to 90% and can be implemented effectively using past data (e.g., "guess what happened") to simulate future forecasting conditions.
  • The primary barrier to widespread forecasting adoption is not a lack of evidence for its utility, but a poor user experience; most people find the initial experience of being wrong and the long feedback loops discouraging without targeted training.
  • OpenAI's Foresight team focuses on measuring macro trends in AI to inform research agendas, policy decisions, and public understanding, distinct from the "Reflection" team (alignment) and "Clarity" team (interpretability).
  • Tom Brown, a current OpenAI researcher, transitioned from an outsider to a full-time employee by securing regular mentorship with co-founder Greg Brockman, illustrating a viable pathway to joining AI safety/capabilities research without a traditional PhD in the field.

AI and Compute Trends

  • Between 2012 and 2018, the amount of compute (floating-point operations) used to train the largest neural networks increased by 10x per year, totaling a 300,000x increase.
  • This exponential growth began after AlexNet (2012), which proved that neural networks could outperform handcrafted heuristics by a significant margin, shifting the field from "research curiosity" to "economic application."
  • The compute growth rate (10x/year) is significantly faster than the pre-2012 trend, which followed a Moore's Law-like curve of roughly 2x every two years.
  • Compute scaling has become the primary driver of immediate capability gains, as hardware can be reallocated and purchased faster than human capital (PhD researchers) can be trained.
  • The widening gap between industrial and academic compute resources risks fragmenting the research ecosystem, suggesting a need for government-provided large-scale compute for academic institutions.
  • Human intelligence appears to be a "scaled-up" version of chimp intelligence (roughly 4x compute), suggesting that massive compute increases can lead to qualitative capability jumps rather than just incremental improvements.

AI Efficiency and Algorithmic Progress

  • Models in 2018 required 25x less compute than 2012 models to achieve AlexNet-level performance, indicating significant algorithmic efficiency gains.
  • This 25x efficiency figure is considered a conservative "floor" because it excludes major breakthroughs like the introduction of new architectures (e.g., Transformers) which may unlock capabilities 100x or 1000x cheaper.
  • Algorithmic progress and compute scaling act multiplicatively rather than additively; a 2x improvement in algorithms combined with a 2x increase in compute yields a 4x total improvement in effective capability.
  • The field of AI efficiency is currently "Goodhart's law" safe because researchers generally optimize for peak performance, not training efficiency, making the measured 25x improvement a reliable metric of actual progress.
  • In specific domains like machine translation, efficiency gains have been even higher (60x in three years), suggesting the 25x figure is an underestimate of total algorithmic progress.
  • "Sample efficiency" (learning from fewer data points) remains a critical, under-met metric; systems that require massive data cannot solve problems where data is scarce or expensive to generate (e.g., industrial safety).

Hardware, Safety, and Career Pathways

  • AI-specific hardware (e.g., GPUs, specialized chips) is a major growth area because general-purpose processors are reaching physical limits, and AI workloads are highly parallelizable.
  • Moore's Law remains the critical long-term variable for long-termists, as its exponent determines the pace of future computational availability, regardless of AI-specific optimizations.
  • Working in AI hardware offers impact through "secure hardware" design to prevent low-level vulnerabilities (like Spectre) and by adopting "Windfall Clauses" (pre-committing excess profits to non-profits).
  • OpenAI's Limited Partnership structure functions as an organizational Windfall Clause, directing all returns above a certain threshold to a non-profit entity.
  • Policy roles regarding AI hardware are currently undersupplied; experts with hardware engineering backgrounds would be highly valuable in government to forecast trends like Moore's Law and inform regulation.
  • A viable career strategy for entering AI research involves securing mentorship by asking potential supervisors specific questions about "what to learn in the next three months," demonstrating focus and reducing the mentor's risk in providing advice.