Interview, Fireside Chat
Ilya Sutskever (OpenAI Chief Scientist) — Why next-token prediction could surpass human intelligence
Alignment and Risks:
- Ilya Sutskever warns against underestimating the difficulty of aligning models smarter than humans, specifically those capable of misrepresenting their intentions.
- He expresses low concern regarding security leaks of model weights, citing strong security measures at OpenAI.
- He anticipates that foreign governments may eventually use open-source models for propaganda or scams, though large-scale tracking of such activities is feasible.
- Sutskever predicts that as models surpass human intelligence, alignment strategies will likely rely on a combination of adversarial stress testing, interpretability via smaller neural nets, and behavioral analysis rather than a single mathematical definition.
AGI Trajectory and Economic Impact:
- Sutskever describes the pre-AGI window as a "good multi-year chunk," noting that AI value grows exponentially, making recent years feel larger than all previous years combined in hindsight.
- He projects OpenAI's 2024 revenue to be around $1 billion, based on extrapolations of growth from GPT-3, the API, and ChatGPT.
- He argues that current "self-driving" level capabilities in AI, while seemingly comprehensive, still lack the reliability required for full AGI, similar to Tesla's current autonomous driving status.
- Reliability is identified as the primary potential bottleneck; if models cannot be trusted to operate without human verification, economic value creation may stall.
- By 2030, Sutskever expects models to be technologically mature and reliable, with a similar trajectory to self-driving cars where the gap between "looks functional" and "actually reliable" closes.
Model Capabilities and Data:
- Sutskever challenges the claim that next-token prediction cannot surpass human performance, arguing that sufficiently smart neural nets can extrapolate the behavior of "hypothetical wise persons" from data.
- He notes that current models are already bad at multi-step reasoning when thinking silently but perform significantly better when allowed to "think out loud."
- While data volumes are currently sufficient, he acknowledges the internet will eventually run out of high-quality training tokens, necessitating alternative training methods.
- Future data sources may include multimodal inputs, but he suggests text-only models can still achieve significant improvements with current data volumes.
- Reinforcement learning data is increasingly generated by AIs themselves, with humans primarily training the reward functions; Sutskever envisions a future where humans perform 1% of the teaching work while AI performs 99%.
- He believes the next paradigm for AGI will likely integrate past ideas rather than being a fundamentally different architecture, though the exact form factor remains uncertain.
Research, Hardware, and Industry Dynamics:
- OpenAI's decision to leave robotics was driven by a lack of viable data pathways at the time; Sutskever states a path forward now exists but requires massive commitment to building and operating thousands of robots.
- He considers Retrieval Transformers promising for storage outside the model but views hardware differences (e.g., NVIDIA GPUs vs. Google TPUs) as minor, with cost per floating-point operation being the primary differentiator.
- Sutskever suggests that while immediate convergence is happening in near-term AI research, divergence will occur in longer-term paths before re-converging once promising directions yield results.
- He argues that AI research will not become a commodity because continuous innovation in reliability, trustworthiness, and model quality will drive demand for new versions.
- Inference costs may rise with model size, but Sutskever posits they will not be prohibitive if the utility provided (e.g., legal advice) justifies the expense, allowing for price discrimination based on model size and quality.
Future Outlook and Human-Machine Interaction:
- Sutskever envisions a post-AGI world where humans become more "enlightened" through interaction with AI, likening it to having the "best meditation teacher in history" available to everyone.
- He speculates that some individuals may choose to become "part AI" to expand their minds and solve complex societal problems.
- Regarding the year 3000, he asserts that no one can predict the state of humanity, but he hopes for a future where humans remain free to evolve morally and solve their own problems with AI as a safety net.
- He believes the AI revolution was somewhat inevitable, likely delayed by only about a year even without specific pioneers, due to the simultaneous maturation of hardware, data availability, and algorithmic understanding.
- He views the "forward-forward" algorithm as a neuroscience-inspired attempt to bypass backpropagation, which is valuable for understanding the brain but unnecessary for engineering high-performance systems.
- Sutskever attributes his career longevity and impact to persistent effort ("trying really hard") and maintaining the right perspective on problems, rather than just being first.
Academic and Corporate Roles:
- He identifies alignment research as a critical area where academic contributions can be highly meaningful.
- While companies currently lead in discovering model capabilities, Sutskever believes academics can still uncover significant insights if they focus on the right problems.
- He argues that the distinction between the "world of bits" and the "world of atoms" is blurring, as AI decisions directly influence physical world actions (e.g., rearranging an apartment).
- Sutskever notes that current AI capabilities have exceeded his 2015 expectations, though he acknowledges his confidence was never 100% at the time.