Conference Presentation, Keynote
Alex Ratner, CEO & Co-Founder, Snorkel AI: Data Centric AI in the Agentic Era
- As AI systems evolve from single-prediction models to complex agentic architectures involving planning, tool use, and multi-turn interactions, the industry expects evaluation of edge cases to become increasingly critical, with well-defined benchmarks (comprising prompts, rubrics, and evaluators) serving as the primary product specifications and sole requirements for building specialized systems.
- Reinforcement learning is predicted to operationalize soon with a "push-button" accessibility model where optimizing data mixtures and having the right ingredients allows users to have models optimized for specific benchmarks, potentially enabling a 7 billion parameter model like Qwen to outperform GPT 4.0 on targeted tasks.
- Building high-quality evaluation benchmarks requires a unified approach between product managers, business stakeholders, and AI teams, as defining "good responses" in specific settings is fundamentally a product management question instantiated into grading rubrics rather than a purely technical challenge.
- Data generation strategies for agentic settings are expected to shift away from relying solely on LLMs to synthetically generate data, which risks a regularization effect that fails to produce novel scenarios, toward injecting human experts into the loop to leverage AI primarily for process acceleration.
- While LLMs acting as judges offer a generic off-the-shelf solution, high-performance outcomes in reinforcement learning are predicted to depend on the development of specialized evaluators and verifiers using smart, programmatic methods, which are expected to yield remarkable results when paired with optimized data sets.
- Agentic systems will be characterized by complex internal processes including tool selection, parallelized execution, summary, and aggregation, or simpler loops driven by significant multi-turn user interaction, all of which will increasingly be defined by the benchmark dataset, grading rubric, and evaluation methods required for their construction.