Interview, Fireside Chat
OpenAI’s IMO Team on Why Models Are Finally Solving Elite-Level Math
- Math model reasoning capabilities have advanced rapidly, evolving from ten-minute sessions to systems capable of sustained concentration for 100 minutes, with aspirations to scale this to thousands or hundreds of thousands of hours to solve major scientific problems.
- The team executed a six-month effort culminating in a two-month sprint involving three core members to target a 2025 IMO gold medal, a goal considered highly unlikely just a year prior.
- Evaluation relied on three external former IMO medalists per proof to achieve unanimous consensus, though the model's outputs were described as beyond the comprehension or grading ability of even the team's expert mathematicians.
- While the internal estimate shifted from less than one-third odds to a strong possibility of success, the team acknowledged models may struggle more than humans on certain distributions, particularly in abstract combinatorics.
- The approach prioritized general-purpose natural language reasoning over formal verification tools like Lean or bespoke systems to maintain agility amidst rapid AI progress.
- The model demonstrated the ability to abstain from answering unsolvable problems like Problem 6 rather than hallucinating solutions, though it performed better on Putnam problems than IMO problems due to differences in time constraints and knowledge requirements.
- Future challenges include shifting from time-boxed competition problems to real research requiring tens of thousands of hours of thinking, such as Millennium Prize problems, with the evaluation process itself becoming a potential bottleneck.
- Current findings are being integrated into broader model updates with plans to deploy improvements to other systems and agents, though widespread access is expected to take time.
- Immediate applications are projected to succeed in written Physics Olympiad sections, while experimental or robotics sections require further development.
- External validation includes ongoing testing by skeptical academics who note the model's improved ability to recognize its own limitations on difficult problems.