Interview, Fireside Chat
Why Voice Will Be the Fundamental Interface for Tech ft ElevenLabs’ Mati Staniszewski
- Origin Story: Inspiration for ElevenLabs emerged in late 2021 when co-founder Piotr (Peter) and the interviewee (Maddie Staniszewski) experienced the monotone, single-voice narration style of Polish-dubbed foreign movies, contrasting it with the native multi-voice audio they remembered from childhood.
- Strategic Focus: The company carved a defensible position by narrowing its focus exclusively on audio, while competitors broadened into multi-modality, avoiding the "roadkill" fate predicted for specialized AI startups.
- Data Architecture: ElevenLabs identified that audio data presents unique hurdles compared to text, specifically the scarcity of high-quality recordings, the frequent lack of accurate transcriptions, and the absence of metadata regarding non-verbal elements like emotion and tone.
- Technical Differentiation: Unlike text models that predict the next token, ElevenLabs' architecture predicts the next sound, requiring the model to understand context (e.g., sarcasm in "what a wonderful day") to adjust emotional delivery and pacing.
- Voice Modeling: The team developed a unique decoding and coding mechanism that allows the model to autonomously determine voice characteristics (gender, age, style) rather than hard-coding specific features, merging text context with a second voice input for high-fidelity synthesis.
- Team Composition: The company employs approximately 15 research and research engineers, sourcing talent globally via a fully remote model, and utilizes a specialized team of "voice coaches" and trained data labelers to manually curate and verify audio emotion data.
- Product Evolution: Early viral moments included book authors using the beta for audiobooks in 2023, followed by the rise of "no-face" narration channels and the release of multi-language dubbing capabilities.
- Enterprise Integration: Major use cases include healthcare automation (e.g., Hippocratic for patient reminders), customer support infrastructure (e.g., Deutsche Telekom), and interactive education (e.g., Chess.com with voices of Magnus Carlsen and Garry Kasparov).
- Agent Deployment: ElevenLabs treats foundation model providers as partners rather than pure competitors, utilizing a cascading mechanism with multiple LLMs to ensure reliability and redundancy in conversational agents.
- Customer Priorities: Enterprise clients prioritize three specific metrics over public benchmarks: expressive quality (language and emotion), low latency for real-time conversation, and infrastructure reliability at scale (e.g., handling millions of simultaneous users like with Epic Games).
- Future Timeline: The company aims to achieve a "Turing Test" pass for human-level voice interaction by 2025, transitioning from current cascading models (speech-to-text-to-speech) to fully duplex "omni" models for true real-time dialogue.
- Safety and Provenance: To combat impersonation, the platform traces all generated audio to specific user accounts and is actively collaborating with academia (e.g., UC Berkeley) to develop open-source detection models for non-AI voice verification.
- European Operations: While based in London with a remote global team, the company cites European access to high-caliber mathematical talent as a primary advantage, though it notes challenges in early-stage mentorship and regulatory uncertainty compared to the US ecosystem.
- Long-term Vision: The interviewee predicts a shift toward "ambient computing" where voice becomes the default interface, enabling universal translation ("Babelfish" effect) and personalized AI tutors that fade technology into the background of human interaction.