newsfilter.io
Conference Presentation, Fireside Chat

RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor

Market Evolution and Mercore's Growth

  • Mercore scaled from a $1B to a $2B revenue run rate within approximately four months.
  • The data market has transitioned from a 2020 crowdsourcing era of low-skilled behavior cloning (SFT and basic RLHF) to a 2024+ "agentic data" era.
  • The new paradigm prioritizes high-skilled expert teams (software engineers, lawyers, doctors, bankers) collaborating to build frontier evaluations and RL environments.
  • Mercore is currently the primary agentic data vendor for all major frontier labs and leading application layer companies (e.g., Harvey, Sierra, Cognition, Ramp).
  • Over the last 24 months, Mercore's expert network throughput grew to 2.5 million hours in Q2 alone, with accelerating growth in expert time used to build environments.

Structure of RL Environments

  • An RL environment consists of three core components:
    • Worlds: Simulated data structures mirroring real-world assets like messages, slides, documents, and sheets.
    • Apps: High-fidelity clones of real software (e.g., Salesforce, ServiceNow, Microsoft 365) accessible via MCP, CLI, or GUI.
    • Tasks: Defined prompts paired with verifiers (rubrics or unit tests) used for both training and evaluation.
  • The primary barrier for frontier labs is covering the full distribution of worlds, apps, and tasks across the entire global economy.
  • Human experts remain essential for creating verifiers because models struggle to self-evaluate their own mistakes in non-simulation domains (e.g., legal or creative tasks).

Case Study: Legal Domain and Performance Gains

  • Mercore published a real-world legal environment built with partners like Latham and Watkins, featuring a full data room and simulated Google Workspace interactions.
  • A specific task evaluated maximum liability for "StarTankers International Limited" vs. "Cooper Jeffries Energy Corporation" under the Oil and Petroleum Act.
  • Rubric criteria were manually constructed by experts to avoid reward hacking and ensure accurate alignment with ground truth.
  • Performance Data: Post-training GLM 4.7 on 1,800 Apex Agent tasks (approx. $500K compute) saw corporate logic scores jump from 4.70% to 26.6%.
  • Generalization: Models showed significant gains on external benchmarks (GPQ Val, Apex V1) despite training only on Mercore-specific data rooms.
  • Current leaderboards now include GLM 5.2 and Kimmy K3, with updated results for K3 pending.
  • Mercore is collaborating with customers like Harvey to build domain-specific environments that enable "frontier intelligence" for vertical applications.

Data Curation and Pricing Models

  • Task-Based Pricing: Custom pricing for complex environments (e.g., $2,000 per task), with some tasks taking humans up to a month to complete; frontier labs may purchase 50,000 tasks monthly.
  • Off-the-Shelf Data: Mercore has invested hundreds of millions into proprietary, reusable datasets sold to multiple customers (e.g., "Neo labs" prefer this model).
  • Hourly Expert Services: A legacy model where customers organize experts themselves, now a smaller focus compared to skilled data offerings.
  • Pricing strategy is backward-looking, calculating the value of reaching a specific leaderboard frontier (e.g., a $1B valuation for a frontier open-source model) against production costs.
  • Cost basis is derived from expert hours (e.g., $150/hour) plus a margin based on the differentiation and frontier nature of the task.
  • Price range for individual tasks spans from $50 to $10,000.

Quality Assurance and Synthetic Data

  • Data quality is measured by Realism (reflecting the real-world distribution of jobs/tasks) and Verifier Accuracy (rubrics scoring trajectories as reliably as humans).
  • Mercore uses "trajectory analysis" (rolling out model trajectories, scoring them, and comparing against human feedback) to calibrate autograders.
  • Synthetic Data Role: RLVR is inherently a bet on synthetic data where models generate trajectories for other models to learn from.
  • Models are used as co-pilots by humans to populate environments and draft tasks, but humans are required to measure capabilities beyond the current model frontier.
  • Cybersecurity is an exception where human grading is less critical because "attacker vs. defender" agents can serve as automated verifiers.

Future Trends and Strategic Advice

  • Ultra-Long Horizon Tasks: The industry is shifting from tasks taking 10 hours to tasks requiring 100 to 1,000 hours of execution.
  • Virtual Co-Workers: A major gap exists in current evaluations: 60-70% of human jobs involve social interaction, yet <1% of evals measure multi-agent social dynamics.
  • AI Co-Pilots in Production: Using AI co-pilots to assist experts in creating tasks and verifiers improves efficiency, though humans remain essential for final rubric definition.
  • Base Model Limitations: A model must show a "pass-at-16" success rate (finding a solution among 16 attempts) to be trainable on complex data; models with 0% success on 16 attempts are generally hopeless for learning.
  • Strategic Recommendation: Companies should leverage partners like Mercore for custom data to maintain competitive moats while utilizing external infrastructure, rather than attempting to build internal talent networks for all data needs.
  • Moat Definition: For application layer companies, data is the most differentiating factor in AI strategy, alongside compute and algorithms.