Conference Presentation, Fireside Chat
RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor
- Dozens of application layer companies are expected to successfully develop industry-leading models with distinct intelligence moats within the next 12 months.
- Updated results for Kimi K3 are anticipated following a post-training run on a 1,800-task dataset previously utilized for GLM 4.7.
- A continuous giant scale-up of diversity across worlds, apps, and tasks is predicted for the going-forward period.
- Development is shifting toward ultra-long horizon tasks aiming for 100 to 1,000 hours of human-level duration, surpassing the current 10-hour limit.
- Virtual co-workers will be introduced to address the realism gap, contrasting the current 1% of evaluations measuring human interaction against the 60-70% of real jobs requiring it.
- The company plans to continue collaborating with customers like Harvey to construct specific domain environments that enable frontier intelligence within their verticals.
- The trend of large-scale apps in 2025 is attributed to the realization that model usefulness is bottlenecked by the ability to utilize context, codebases, and laptop tools.
- A primary barrier for frontier labs is identified as the necessity to cover the full distribution of all worlds, apps, and tasks in the economy.
- Most domains currently face significant challenges in reliably identifying model mistakes without human-created rubrics.
- Building in-house talent networks and infrastructure is described as difficult compared to leveraging partners with economies of scale, though some companies attempt the former.
- Custom task purchasing may reach dramatic scales, with certain frontier labs potentially buying 50,000 tasks per month.
- Human completion times for customer tasks generally vary from a few hours up to one month depending on the specific task.
- Data pricing is forecasted to range from $50 to $10,000 per task.
- The business focus is shifting away from hourly expert services toward skilled data offerings over time.
- K3 is identified as a rare exception capable of creating tasks that inferior models can learn from, enabling the distillation of human task creation.