Interview
John Schulman (OpenAI Cofounder) — Reasoning, RLHF, & plan for 2027 AGI
- Over the next one to two years, models are expected to transition from suggesting single functions to executing whole coding projects, including writing multiple files, testing output, and iterating on instructions.
- Progress toward AGI is considered possible within the next year if no new bottlenecks emerge, though the speaker emphasizes a preference for incremental capability releases over discontinuous jumps.
- Significant model improvements are predicted over a five-year horizon, driven by training on long-horizon tasks which may reveal phase transitions where models handle much longer sequences more efficiently.
- Future models are anticipated to use a unified planning mechanism for various time scales (e.g., one month to 100 years) and improve sample efficiency through better error recovery and edge case handling.
- Training data requirements for longer tasks may increase, and scaling laws for these specific tasks might not be clean unless experiments are carefully designed.
- Current limitations regarding user interface interaction, such as using websites designed for humans, are expected to be resolved via vision capabilities, potentially reducing the immediate need for API-driven web redesigns.
- Fine-tuning with English text data is expected to generalize to other languages and modalities, with small datasets (e.g., ~30 examples) capable of teaching a model to explain its limitations effectively across untrained capabilities.
- If AGI arrives sooner than expected, the outlook suggests pausing training and deployment to ensure safety, requiring coordination among entities to prevent unsafe race dynamics.
- Strategic safety measures include sandboxing, robust monitoring for immediate failure detection, and a defense-in-depth strategy combining aligned models with external safeguards.
- The timeline for job replacement is estimated at five years, with a continued reliance on human oversight in firm operations to mitigate tail risks associated with fully autonomous AI-run entities.
- AI is expected to increasingly assist in sophisticated tasks like research and coding, potentially accelerating scientific advancement by processing literature and data beyond human patience limits.
- A shift in training paradigms is anticipated toward combining training-time computation with test-time reasoning (in-context learning), addressing the missing middle ground of medium-term memory and active learning.
- Post-training methods, such as RLHF and the integration of instruction and chat data, have driven significant capability gains, including reduced hallucinations and improved instruction following compared to pre-trained models.
- The field may face challenges related to data limits in the future, potentially shifting the nature of pre-training and encouraging larger models to learn better shared representations rather than memorization.
- Deployment strategies aim to balance user helpfulness with broader societal constraints, avoiding the imposition of specific opinions while blocking usage that harms others.
- Industry consolidation risks are noted, though smaller players may catch up through distillation techniques that clone outputs from larger models.
- Human labor in model training remains critical, with non-expert labelers often performing as well as researchers due to the base model's existing knowledge and generalization capabilities.
- The form factor of AI interaction is trending toward proactive, collaborative assistants integrated into daily work, moving away from simple query-response models to background project assistance.