newsfilter.io

Latest Interviews

Showing 1–1 of 1 transcripts.

Clear all filters
  1. a16z1h 45m

    Beyond Leaderboards: LMArena’s Mission to Make AI Reliable

    Anjney Midha, Anastasios N. Angelopoulos, Wei-Lin Chiang, Ion Stoica

    LM Arena has transformed from a static benchmark into a dynamic "humanity's exam" that evaluates over 280 AI models through real-time feedback from one million monthly users, effectively eliminating data contamination through fresh prompt generation. By treating evaluation as Reinforcement Learning rather than Supervised Learning, the platform utilizes techniques like "style control" and the open-sourced "Prompt-to-Leaderboard" router to achieve twice the performance-per-cost while maintaining academic neutrality. Looking forward, the organization plans to expand into private industry-specific Arenas and multi-modal agent testing while remaining committed to open-sourcing all data and research to preserve ecosystem trust.