newsfilter.io
Fireside Chat, Interview

From Vibe Coding to Vibe Researching: OpenAI’s Mark Chen and Jakub Pachocki

  • The organization aims to automate the discovery of new ideas, machine learning research, and progress in other sciences, moving beyond current capabilities where models operate for one to five hours to handle tasks spanning months and years within the next one to five years.
  • GPT-5 and subsequent models are positioned as foundational for bringing reasoning and agentic behavior to the mainstream, with a focus on solving open-ended problems over longer time horizons that currently require human persistence, such as proving millennium prize problems or developing entire programs.
  • Evaluation priorities are shifting from saturated math and programming competition metrics toward economically relevant movements and "vibe researching," addressing a current deficit in great evaluations while anticipating that the gap in solving the hardest programming problems will close rapidly.
  • Over the next two years, the organization expects reward modeling and fine-tuning processes to evolve rapidly, becoming simpler while inching toward more human-like learning, alongside improvements in model memory, planning capabilities, and quality retention over long horizons.
  • The Codex team is developing reasoning models for real-world coding that adapt latency based on problem difficulty—prioritizing speed for easy tasks and higher latency for difficult ones to ensure optimal solutions—and aims to exit the "uncanny valley" of AI coding tools by handling messy environments and defining behavioral presets.
  • Strategic plans include a dedicated group for algorithmic advances, a potential future focus on robotics, and a flexible resource allocation strategy regarding compute, data curation, and personnel, with the belief that a data-constrained regime is not imminent in the near term.
  • The organization anticipates that deep learning will continue to discover ideas that work more often than in the past and expects to avoid the learning plateaus seen at other companies, driven by a strong belief in its long-term research program independent of short-term product reception.
  • Risks and challenges include the difficulty of maintaining quality and consistency over multi-year horizons, the complexity of extending reasoning to domains that are less verifiable than well-posed constraint problems, and the danger of failing to achieve leadership in all key areas if the organization becomes "second place at everything."