Fireside Chat, Interview
How Intelligent Is AI, Really?
- The ARC Prize Foundation plans to advance progress toward systems capable of human-like generalization by shifting focus to long-horizon learning that spans from hours to a lifetime, prioritizing efficiency in acquiring new knowledge over specific environment mastery.
- Arc AGI 3 is scheduled for release next year as a completely new interactive benchmark featuring approximately 150 video game environments that contain no instructions, language, or symbols.
- To ensure accessibility and relevance, the foundation will recruit members of the general public, such as accountants and Uber drivers, to validate the benchmark; any game failing to be solved by 10 specific individuals will be excluded from the final set.
- Performance metrics in Arc AGI 3 will measure intelligence by the efficiency of actions taken, comparing the number of steps an AI requires against human norms rather than relying solely on accuracy.
- AI performance will be normalized against average human test results, and any system achieving 100% accuracy will be analyzed to identify remaining failure points as the foundation prepares to declare the achievement of AGI.
- While solving Arc AGI 3 may not be sufficient to declare AGI, it is expected to provide the most authoritative evidence to date regarding a system's generalization capabilities, with investment directed toward such generalizing systems rather than those requiring specific training environments.
- The foundation will maintain a hidden test set for future benchmarks and intends to continue testing all tasks to ensure normal people can complete them, contrasting with benchmarks designed to solve "PhD plus plus problems."
- The organization expects the community to recognize reasoning as a transformational shift for AI, viewing big lab endorsements as secondary to its primary mission of inspiring researchers, small teams, and individuals.