Interview, Fireside Chat, Keynote
How Far Are We From An AI Einstein? - Adam Brown
Future Trajectory of AI Capabilities
- The speaker posits that a terminal milestone for Large Language Models (LLMs) would be the ability to derive General Relativity from Newtonian physics using current laws.
- Achieving this feat is estimated to occur within approximately 10 years, at which point AI would fully encompass human intelligence.
- Once this threshold is crossed, the speaker suggests there may be little remaining for humans to achieve in terms of intellectual discovery.
Nature of Intelligence and Abstraction
- LLMs appear to function as interpolators, but the level of abstraction they operate on continues to rise significantly.
- The speaker hypothesizes that from a sufficiently high level of abstraction, the invention of General Relativity is merely interpolation, potentially mirroring human intelligence mechanisms.
- While LLMs possess billions of parameters akin to the human brain, they do not inherently "think" in high-dimensional spatial spaces any more than humans do.
- Humans have historically relied on notation (e.g., tensor notation, Einstein summation convention) to manipulate high-dimensional concepts rather than intuitively perceiving them.
AI vs. Human Representation Learning
- AI may develop more sophisticated representations of complex geometries by processing vastly larger datasets of problems than any human could encounter.
- The historical power of physics breakthroughs is attributed to the invention of new notations and representations (e.g., Penrose's work), suggesting AI could similarly contribute to physics via novel representational frameworks.
- Unlike humans, LLMs may not require human-compatible representations to utilize their knowledge effectively.
Knowledge Translation and Discovery
- Despite LLMs' overwhelming advantage in knowledge retention (reading more than any human in a lifetime), they currently struggle to translate this into novel discoveries or "conceptual leaps."
- The speaker compares this limitation to chess engines: while they search vastly more positions than humans, their ability to evaluate positions is less "natural," implying similar gaps in reasoning versus raw calculation for LLMs in physics.
- A hypothetical human who memorized LLM-level knowledge and open problems across all fields would likely still lack the intuitive leap capabilities of Einstein, though they might make basic correlation-based discoveries (e.g., magnesium and headaches).
Performance on Graduate-Level Exams
- Private evaluations by a Stanford professor reveal a dramatic improvement in LLM performance on his graduate General Relativity exams over a three-year period.
- Three years ago, models scored zero; one year ago, they performed at a weak student level; currently, they "essentially ace" the final exam.
- The professor notes he is retiring the specific exam used for testing because the models now pass it too easily.
- Success on these exams requires two distinct capabilities: translating word-based physics problems into mathematical formulations and solving the resulting mathematics.
Trends in Evaluation Difficulty
- The difficulty of evaluating LLMs has increased exponentially as the models have advanced.
- Until recently, scraping standard high school math problems was sufficient to test LLMs; today, researchers must commission unique, PhD-level problems to challenge current systems.
- It remains unclear if models can solve hard research problems; current tests suggest a gap between passing exams and solving novel research challenges.