Interview, Fireside Chat
Francois Chollet — Why the biggest AI models can't solve simple puzzles
- The $500,000 prize for achieving 85% on the ARC benchmark will be withdrawn if the system "survives three months," with the full contest running through mid-November and the $100,000 progress prize awarded at the end of November.
- Organizers plan to use the downtime between December and February to re-baseline the community based on top scores before resuming the contest next year, with the ultimate goal of running until 85% is achieved.
- "ARC 2" is scheduled for release later this year to address dataset redundancy and flaws, while the private test set will eventually be accessed via API to prevent data leakage common with the current public GitHub dataset.
- Progress toward AGI is predicted to be set back by "5 to 10 years" due to the closure of frontier research publishing and the redirection of resources toward LLM variations, which are seen as sucking the oxygen out of other AI approaches.
- Skepticism exists that LLMs will reach 80% on the ARC benchmark within a year, with the belief that true AGI requires the ability to adapt to novelty on the fly rather than relying on brute-forcing training data or interpolation.
- The anticipated path to solving ARC and achieving AGI involves a hybrid system combining deep learning for "intuition" to guide discrete program search, avoiding the "combinatorial explosion" of brute force or "shallow recombination" methods.
- Future benchmark versions may utilize test server APIs to ensure models cannot be trained on private test data, addressing the risk of "hacking" through methodological loopholes or data leakage.
- The contest is expected to reveal if the benchmark is resistant to "brute-force attempts involving millions of training puzzles," with the possibility that increasing the prize significantly might quickly expose low-hanging fruit if such opportunities exist.
- Multimodal models native to spatial reasoning are expected to outperform current text-based models within a few months, though scaling up compute alone is viewed as insufficient without architectural changes to handle "system 2" thinking like explicit planning.
- While LLMs may automate tasks in static distributions, intelligence is defined as the capacity to handle "novelty" and uncertainty, with predictions that software engineering jobs will increase over the next five years due to the dynamic nature of the field.
- Human children acquire core knowledge, such as basic physics and object permanence, within the first three to four years of life, a capability that contrasts with the "memorization plus regularization" approach of current models like GPT-4 or Gemini 1.5.
- The speakers expect that solutions to ARC will likely need to leverage aspects of deep learning and LLMs while integrating discrete program search, representing a middle ground between different computational paradigms.
- Open-source ecosystems are regarded as the most powerful innovation drivers, with concerns that the trend toward closed research by major labs hinders the generation of new ideas necessary to break through current plateaus.
- If a closed-source model were to achieve over 50% on the private test set, it could shift the organizers' view on whether increased compute is the solution, whereas a brute-force solution would suggest intelligence is merely getting to the right part of a distribution.
- The speakers anticipate that "GPT-5" or similar models will not alter their perspective on AGI unless they demonstrate genuine adaptation to novelty, distinguishing true generalization from the "interpolation" capabilities of current parameter curves.
- The "Jack Cole approach" utilizing test-time fine-tuning is viewed as a form of program synthesis that addresses the lack of active inference, though it may rely on synthetic data due to GPU constraints like P100 limits.
- Human intelligence is characterized by efficient neural wiring and the ability to synthesize programs from very few examples, whereas current deep learning models require dense sampling and struggle with extreme generalization without architectural shifts.
- The current competition serves as a stress test to determine if the benchmark can be "hacked" by large models given API access, which could reveal if frontier models like OpenAI are training on such calls.
- Automation via memorization is deemed sufficient only for predictable environments, while tasks involving constant novelty, such as software development, require the "pathfinding" capabilities of true intelligence.
- The organizers expect that increasing the prize will attract more participants to rigorously test if the benchmark is hackable, ensuring that any winning solution represents a "huge milestone" rather than a circumvention of the task's intent.