Fireside Chat, Interview
ElevenLabs' Mati Staniszewski: How Voice Becomes the Interface for AI
Origins and Human Context
- Founding Timeline: Eleven Labs was established in 2022 by co-founders Gabriel Sanchez and Piotr, who met as best friends in high school in the suburbs of Warsaw, Poland.
- Market Gap Identification: The founders identified a specific deficiency in Polish audio localization, where foreign movies utilize a single monotone narrator for all characters, stripping content of emotion and intonation.
- Strategic Vision: This observation drove the company's mission to enable universal audio communication where any individual can speak any language with authentic emotional nuance and intonation.
- Core Belief: The founders posit that as robotics and AI integrate into daily life, voice will become the primary interface for human-machine interaction.
Strategic Approach and Execution
- Counter-Cyclical Timing: The company launched in 2022, a period dominated by crypto and metaverse hype when audio AI was considered a niche with minimal research focus.
- Computational Efficiency: Unlike peers requiring hundreds of billions in compute, Eleven Labs leveraged the fact that early audio models required significantly less computational power than text or visual models.
- Remote-First Talent Acquisition: The company adopted a remote-first strategy connecting researchers in London and Warsaw, recruiting based on GitHub contributions and technical work rather than geographic proximity.
- Rapid Monetization: The company prioritized immediate revenue generation to fund model development, aiming to maintain healthy margins and financial independence before scaling external capital.
- Current Scale (Q1): The company generated over $100 million in net new Annual Recurring Revenue (ARR) in Q1.
- Workforce Size: As of the discussion, the company employs over 400 people and generates over $400 million in revenue while maintaining small, flat teams.
Product Evolution and Model Architecture
- Foundational Models: The initial product stack consisted of three core components: text-to-speech (understanding context/emotion), speech-to-text (transcription), and localization (translation + synthesis).
- Real-Time Integration: The product line expanded to include real-time streaming models and conversational orchestration to build voice agents capable of turn-taking.
- Musical Capabilities: Eleven Labs developed audio synthesis models for music generation, noted as a significantly harder modality to master than speech.
- Emotional Intelligence: The current R&D focus is on "emotional intelligence," enabling voice agents to detect user stress or excitement and adjust their tone, pacing, and reassurance accordingly.
- Audio General Intelligence: Future roadmap items include combining modalities (e.g., narrating, pausing, then singing) within a single continuous audio stream.
Market Applications and Use Cases
- Customer Support: The most widespread application involves replacing traditional phone trees with voice agents for support inquiries.
- Revenue Generation: Voice agents are increasingly deployed for sales, including outbound operations (e.g., Deliveroo agents updating restaurant hours) and inbound lead qualification.
- Government Services: The Ukrainian government utilizes Eleven Labs voice agents to provide citizens with real-time information regarding war conditions, safety protocols, and educational content.
- Interactive Education: Platforms like MasterClass use voice agents to create dynamic, 24/7 interactive learning environments where users can negotiate or practice skills with AI avatars of experts (e.g., Gordon Ramsay, Chris Voss).
- Agent-to-Agent Communication: Early tests show agents can communicate with each other more efficiently than humans, potentially shifting to non-verbal or encoded data streams for speed.
Organizational Structure and Operations
- Small Unit Model: Teams remain small, averaging fewer than 10 people per research or product unit, including non-technical functions like legal and Go-To-Market.
- Embedded Engineering: Non-technical teams include dedicated engineers to automate workflows, upskill colleagues via "vibe coding," and ensure security compliance.
- Title Elimination: The company operates without formal job titles to optimize for impact and allow rapid internal growth based on contribution rather than tenure.
- Automated Negotiation Scoring: A scoring system was introduced to automate sales contract terms (e.g., indemnity caps), allowing the system to grant specific concessions based on customer size without human intervention.
Future Outlook and Challenges
- Trust and Authentication: As voice AI becomes ubiquitous, the market will shift from detecting AI to verifying "real," authenticated humans, necessitating robust watermarking and encryption.
- Jagged Intelligence: Significant performance gaps remain in deep emotional interactions (agents still struggle with nuanced emotion delivery) and generating top-tier commercial music.
- Training Philosophy: The company prioritizes long-term impact over short-term revenue, investing heavily in data annotation of the "how" of audio (emotion, style) by specialists like voice coaches.
- Defensibility: Competitive advantages are derived from proprietary data collection (user preferences), domain-specific fine-tuning, and the ecosystem of over 20,000 user-contributed voices.
- Robotics Interface: The company anticipates voice becoming the critical bottleneck and unlock for interacting with future robotic intelligence.