Latest Interviews
Showing 1–2 of 2 transcripts.
Clear all filters- a16z1h 45m
Beyond Leaderboards: LMArena’s Mission to Make AI Reliable
Anjney Midha, Anastasios N. Angelopoulos, Wei-Lin Chiang, Ion Stoica
LM Arena has transformed from a static benchmark into a dynamic "humanity's exam" that evaluates over 280 AI models through real-time feedback from one million monthly users, effectively eliminating data contamination through fresh prompt generation. By treating evaluation as Reinforcement Learning rather than Supervised Learning, the platform utilizes techniques like "style control" and the open-sourced "Prompt-to-Leaderboard" router to achieve twice the performance-per-cost while maintaining academic neutrality. Looking forward, the organization plans to expand into private industry-specific Arenas and multi-modal agent testing while remaining committed to open-sourcing all data and research to preserve ecosystem trust.
- a16z1h 36m
Building the Next Generation of Conversational AI
Ankit Kumar, Anjney Midha, Maya
Sesame is developing a voice-first "companion" interface using a talent-dense team of fewer than fifteen engineers to prioritize natural conversational dynamics over general-purpose utility. The company has open-sourced its Conversational Speech Model base weights while withholding character-specific implementations, aiming to evolve toward a full duplex architecture capable of native audio understanding and real-time interruption handling. By targeting smart glasses as the optimal hardware form factor and employing qualitative human evaluation rather than standard metrics, Sesame seeks to build a long-term memory layer that acts as an emotionally resonant mediator for multi-step tasks.