newsfilter.io
Conference Presentation, Fireside Chat

RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

  • Dozens of application layer companies are expected to successfully develop industry-leading models with distinct intelligence moats within the next 12 months.
  • Updated results for Kimi K3 are anticipated following a post-training run on a 1,800-task dataset previously utilized for GLM 4.7.
  • A continuous giant scale-up of diversity across worlds, apps, and tasks is predicted for the going-forward period.
  • Development is shifting toward ultra-long horizon tasks aiming for 100 to 1,000 hours of human-level duration, surpassing the current 10-hour limit.
  • Virtual co-workers will be introduced to address the realism gap, contrasting the current 1% of evaluations measuring human interaction against the 60-70% of real jobs requiring it.
  • The company plans to continue collaborating with customers like Harvey to construct specific domain environments that enable frontier intelligence within their verticals.
  • The trend of large-scale apps in 2025 is attributed to the realization that model usefulness is bottlenecked by the ability to utilize context, codebases, and laptop tools.
  • A primary barrier for frontier labs is identified as the necessity to cover the full distribution of all worlds, apps, and tasks in the economy.
  • Most domains currently face significant challenges in reliably identifying model mistakes without human-created rubrics.
  • Building in-house talent networks and infrastructure is described as difficult compared to leveraging partners with economies of scale, though some companies attempt the former.
  • Custom task purchasing may reach dramatic scales, with certain frontier labs potentially buying 50,000 tasks per month.
  • Human completion times for customer tasks generally vary from a few hours up to one month depending on the specific task.
  • Data pricing is forecasted to range from $50 to $10,000 per task.
  • The business focus is shifting away from hourly expert services toward skilled data offerings over time.
  • K3 is identified as a rare exception capable of creating tasks that inferior models can learn from, enabling the distillation of human task creation.