Fireside Chat, Interview
From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki
- Strategic Goal: OpenAI's primary research objective is to produce an "automated researcher" capable of independently discovering new ideas and automating progress in ML and other scientific domains.
- GPT-5 Launch Thesis: The GPT-5 release aims to mainstream "reasoning" by eliminating the user's need to choose between instant-response models (GPT-3/4) and deep-thought models (O-series), delivering reasoning and agentic behavior by default.
- Evaluations and Milestones:
- Current benchmarks (e.g., AtCoder, IOI, IMO) are becoming saturated, with incremental gains (e.g., 96% to 98%) considered less significant.
- Future success metrics will focus on the model's ability to "discover new things" and generate economically relevant movement rather than just solving contest problems.
- Models have achieved runner-up performance in the AtCoder competition, with plans to target number one.
- Reasoning and Time Horizons:
- GPT-5 demonstrates near-mastery of high school-level competition problems, extending reasoning horizons to approximately one to five hours.
- Future development will focus on extending these horizons, improving long-term planning, and retaining memory over extended autonomous operations.
- Reasoning is identified as the core capability required for agents to operate robustly over long periods, allowing for error correction and strategic pivots.
- Domain Application:
- Progress in reasoning is expected to transfer from hard sciences (math, physics) to open-ended research where "right/wrong" answers are less explicit, such as formulating research hypotheses.
- Internal testing with professional physicists and mathematicians revealed GPT-5 can automate tasks that previously took students months, described as a "lightbulb moment" for domain experts.
- Reinforcement Learning (RL):
- RL is viewed as a versatile, non-plateauing method for aligning models to the real world, particularly after the language modeling breakthrough provided a rich environment.
- The industry is moving toward human-like learning processes, suggesting that current reward modeling techniques will evolve to become simpler over the next two years.
- GPT-5 Codex and Coding:
- New coding models leverage the intelligence of reasoning models to handle messy, real-world environments and define behavioral "specifications" (e.g., latency vs. solution quality trade-offs).
- High school students report that "vibe coding" (generative, intuitive coding) is becoming the default workflow, moving away from writing code manually from scratch.
- Mark and Jakob acknowledge that current coding models have surpassed their own competitive programming capabilities, though they note an "uncanny valley" remains before models act as fully competent co-workers.
- Research Culture and Talent:
- OpenAI retains top talent by focusing on fundamental research rather than copying competitors, fostering a culture of innovating at the frontier.
- Hiring criteria prioritize individuals who have solved hard problems in diverse fields and possess the ability to persist through failure while remaining truth-seeking.
- The organization distinguishes between two researcher archetypes: those strong in idea generation and those strong in rigorous experimental execution, valuing both styles.
- Organizational Structure and Priorities:
- A clear separation exists between product teams accountable for product success and fundamental research teams protected from short-term market pressures.
- Compute remains the primary constraint; the industry is not expected to shift to a "data-constrained" regime soon.
- Long-term research roadmaps are driven by strong convictions (e.g., automated research) rather than short-term product reception, though product feedback informs iterative strategy.
- Leadership Dynamics:
- The partnership between Mark Sutskever (Chief Scientist) and Jakob Sutskever (Chief Research Officer) is built on deep trust developed during early work on reasoning models, characterized by mutual respect for technical depth and organizational leadership.
- This synergy allows for the coordination of diverse research bets (e.g., diffusion, reasoning, agents) into a coherent long-term vision without stifling bottom-up exploration.
- Forward-Looking Statements:
- Within the next year, models are expected to solve problems with significantly longer time horizons than the current one-to-five-hour window.
- Robotics and physical constraints (energy, hardware) are identified as the next major frontiers for AI application beyond pure intelligence.
- The "uncanny valley" in coding tools is expected to be resolved soon, making AI assistance the standard for complex refactorings (e.g., 30-file changes in 15 minutes).