newsfilter.io
Interview, Fireside Chat, Keynote

How Far Are We From An AI Einstein? - Adam Brown

  • Future Trajectory of AI Capabilities

    • The speaker posits that a terminal milestone for Large Language Models (LLMs) would be the ability to derive General Relativity from Newtonian physics using current laws.
    • Achieving this feat is estimated to occur within approximately 10 years, at which point AI would fully encompass human intelligence.
    • Once this threshold is crossed, the speaker suggests there may be little remaining for humans to achieve in terms of intellectual discovery.
  • Nature of Intelligence and Abstraction

    • LLMs appear to function as interpolators, but the level of abstraction they operate on continues to rise significantly.
    • The speaker hypothesizes that from a sufficiently high level of abstraction, the invention of General Relativity is merely interpolation, potentially mirroring human intelligence mechanisms.
    • While LLMs possess billions of parameters akin to the human brain, they do not inherently "think" in high-dimensional spatial spaces any more than humans do.
    • Humans have historically relied on notation (e.g., tensor notation, Einstein summation convention) to manipulate high-dimensional concepts rather than intuitively perceiving them.
  • AI vs. Human Representation Learning

    • AI may develop more sophisticated representations of complex geometries by processing vastly larger datasets of problems than any human could encounter.
    • The historical power of physics breakthroughs is attributed to the invention of new notations and representations (e.g., Penrose's work), suggesting AI could similarly contribute to physics via novel representational frameworks.
    • Unlike humans, LLMs may not require human-compatible representations to utilize their knowledge effectively.
  • Knowledge Translation and Discovery

    • Despite LLMs' overwhelming advantage in knowledge retention (reading more than any human in a lifetime), they currently struggle to translate this into novel discoveries or "conceptual leaps."
    • The speaker compares this limitation to chess engines: while they search vastly more positions than humans, their ability to evaluate positions is less "natural," implying similar gaps in reasoning versus raw calculation for LLMs in physics.
    • A hypothetical human who memorized LLM-level knowledge and open problems across all fields would likely still lack the intuitive leap capabilities of Einstein, though they might make basic correlation-based discoveries (e.g., magnesium and headaches).
  • Performance on Graduate-Level Exams

    • Private evaluations by a Stanford professor reveal a dramatic improvement in LLM performance on his graduate General Relativity exams over a three-year period.
    • Three years ago, models scored zero; one year ago, they performed at a weak student level; currently, they "essentially ace" the final exam.
    • The professor notes he is retiring the specific exam used for testing because the models now pass it too easily.
    • Success on these exams requires two distinct capabilities: translating word-based physics problems into mathematical formulations and solving the resulting mathematics.
  • Trends in Evaluation Difficulty

    • The difficulty of evaluating LLMs has increased exponentially as the models have advanced.
    • Until recently, scraping standard high school math problems was sufficient to test LLMs; today, researchers must commission unique, PhD-level problems to challenge current systems.
    • It remains unclear if models can solve hard research problems; current tests suggest a gap between passing exams and solving novel research challenges.