newsfilter.io
Conference Presentation, Keynote

Agents are Here. Compute has to Change | Sachin Katti, OpenAI | RAISE Summit 2026

  • OpenAI's internal roadmap, originally published years ago, accurately predicted the industry's progression from chatbots (2022) to reasoning models (2024, O-series) to the current "agentic era."
  • The organization is currently in the "takeoff" phase of agent adoption, utilizing AI to define, execute, and help conduct its own next-generation research.
  • OpenAI is on the cusp of deploying "AI interns," where AI systems are already performing significant portions of internal research and AI engineering tasks.
  • Full recursion, where AI builds the next generation of AI, remains a consistent long-term prediction within the company's strategic framework.
  • Internal data indicates that nearly 100% of OpenAI's output tokens are now generated by Codex, marking a total shift to an agentic interface for daily operations.
  • This agentic adoption is not limited to engineers; non-engineering functions (customer support, legal, sales, go-to-market) are driving substantial token consumption growth.
  • Research-side token usage has grown nearly 50x over the last few months, driven by AI-assisted experimentation.
  • The acceleration of AI-assisted research has compressed model release cycles from 6–9 months to approximately one new model per month.
  • Scaling laws for loss continue to hold for pre-training, negating the perception that pre-training has slowed down.
  • Future research compute demands are expanding beyond pre-training to include synthetic data, post-training, and test-time compute with Reinforcement Learning (RL).
  • AI agents enable researchers to run many more experiments in parallel, causing the total volume of required research compute to explode rather than just the size of individual models.
  • Agentic product workloads differ fundamentally from chatbots by involving planning, tool calling, iterative observation, and non-linear task execution.
  • Agents accumulate significant context over longer horizons and perform tool interleaving, requiring diverse compute resources beyond simple GPU inference.
  • User feedback emphasizes a desire to remain in a state of "flow," necessitating shorter tool response times to prevent context switching during complex tasks.
  • Product compute infrastructure is becoming heterogeneous, integrating GPU clusters with extensive CPU and storage resources to handle rich context and varied latency profiles.
  • Future compute needs will require complex networking to connect diverse system components to support instant responses, parallel processing, and deep research tasks.
  • The pace of model releases is expected to accelerate further as AI-driven research loops continue to reduce iteration times and increase experiment volume.