newsfilter.io

Stanford CS153 Frontier Systems | Mati Staniszewski from ElevenLabs on The Future of Voice Systems

ElevenLabs Origins and Foundational Philosophy

  • ElevenLabs was founded in London (with co-founders Matt Garman and Piotr, both former Google/Palantir employees) after a realization that Polish movie dubbing suffered from "Frankenstein" voice quality where one narrator read all characters with a flat tone.
  • The company initially ran its internal operations on Discord to avoid meetings and emails, treating the platform as a "petri dish" for rapid community iteration before eventually migrating to Slack.
  • The initial product strategy focused on a Product-Led Growth (PLG) motion, engaging directly with creators and developers to validate if text-to-speech quality was sufficient for real-world use cases.
  • The founding team identified two distinct problems to solve: fixing foundational research for audio/voice and applying those models to specific customer pain points like voiceover corrections and dubbing.
  • ElevenLabs chose to isolate and perfect the "last mile" (text-to-speech) component of the intelligence pipeline first, rather than building a full-stack AI dubbing system immediately, as the transcription and translation models were not yet reliable enough in 2022.
  • The early technical approach leveraged the open-source "Tortoise" model by James Betker, which offered high-quality prosody but suffered from extremely slow generation speeds and instability over long sequences.
  • In the earliest stages, the team self-funded operations, spending tens of thousands (approaching $100k) on compute, while explicitly rejecting the $6,000 cost of patent filing to prioritize rapid innovation over defensive IP protection.

Evolution of the AI Audio Stack (2022–2026)

  • 2022: The team achieved the first breakthrough in context-aware, emotional text-to-speech generation, moving away from hard-coded voice parameters (age, accent) to abstracted model definitions that replicated natural delivery.
  • 2023: The focus expanded to wider voiceover, narration, and voice cloning, introducing a marketplace for voice contributions and building tooling for audiobooks and video voiceovers.
  • 2024: The arrival of high-quality transcription models combined with LLMs and speech generation enabled true "AI dubbing," demonstrated by the ability to dub Javier Milei's UN speech (translating from Spanish to English while preserving his unique tone) and conversations between world leaders like Zelensky and Modi.
  • 2025: The technology evolved to support real-time, interactive voice agents capable of detecting user emotion (stress, excitement) and responding with matching tonality, closing the loop between Speech-to-Text, LLM reasoning, and expressive Text-to-Speech.
  • 2026 (Projected): The industry is expected to shift from strictly "cascaded" architectures (separate models for transcription, translation, and generation) toward "fused" omni-models that combine modalities to reduce latency, though cascaded systems will likely remain preferred for high-reliability enterprise tasks.
  • ElevenLabs predicts that emotionality can be fixed in both cascaded and fused architectures by explicitly labeling sentiment data and passing emotional context as parameters to the generation model.

Business Performance, Growth, and Organizational Structure

  • ElevenLabs reached over $330M in Annual Recurring Revenue (ARR) by the end of 2025, with the most recent quarter adding over $100M, bringing total ARR to approximately $430M within a 36-month period.
  • The company employs over 450 people across London, New York, Warsaw, and San Francisco, operating with a "small team" philosophy where squads are kept under 10 members to maximize ownership and decision-making speed.
  • Revenue growth is driven by a hybrid model: over 50% comes from enterprise sales (forward-deployed engineering with clients like Deutsche Telekom, Revolut, and Klarna) and roughly 50% from PLG/self-serve channels.
  • The pricing strategy is strictly value-based, aiming to capture roughly 10% of the economic value delivered to the customer, rather than pricing based on computational costs or token usage.
  • Predictability in enterprise revenue is achieved through "forward-deployed engineering" teams that deeply integrate with clients to solve specific business problems, whereas PLG growth remains less predictable but offers broader innovation access to individual creators.

Strategic Partnerships, Safety, and Market Dynamics

  • Matt Garman emphasizes a collaborative industry stance, noting that competitors like Sesame (founded by former Oculus CEO Brendan) are essential partners; Sesame CEO Brendan has also angel-invested in ElevenLabs, and ElevenLabs has supported Sesame's open-source CSM model.
  • Safety and security are addressed by baking provenance and watermarking directly into the models, allowing for the tracing of generated content to specific users and preventing fraud or deepfake scams.
  • The company advises against using voice alone for authentication in banking or security systems due to the ease of replication, instead promoting multi-factor authentication or using voice agents to counter-scammers by wasting their time.
  • ElevenLabs has successfully synthesized voices for nearly 10,000 people who have lost their voice due to ALS or throat cancer, demonstrating the technology's humanitarian potential.
  • The company is supporting the Ukrainian government through the "DIA" app and government ministries, providing voice access to citizen services during the war, while maintaining a strict Western-aligned policy on technology transfer.
  • Regarding global competition, ElevenLabs acknowledges strong open-source capabilities from China (e.g., in video models like Seedance) but differentiates through a focus on open ecosystem participation, safety watermarking, and the "middle-to-middle" AI workflow rather than end-to-end automation.

Future Roadmap and Application Scenarios

  • On-Device Models: ElevenLabs has successfully constrained its text-to-speech models to run on-device, prioritizing quality before deployment, though current on-device versions lack the full interactivity and emotional transfer capabilities of cloud versions.
  • The "Five-Year" Vision: The company aims to become one of the three to five dominant "conversational platforms" (similar to the cloud stack) that handle all interactions between businesses and audiences, from sales and support to marketing and internal training.
  • Studio Adoption: Major production studios are shifting from skepticism to adoption as "middle-to-middle" tools emerge, allowing directors to control emotional delivery (e.g., "more dramatic," "slower") rather than relying on the model's autonomous interpretation.
  • Future Bottlenecks: The primary research focus remains on combining modalities to create truly personalized, real-time agents that understand user preferences and can pull from enterprise knowledge bases without hallucinating.
  • Economic Models: The industry is currently working to resolve the economic models for AI voiceovers, specifically regarding how to compensate voice actors whose work is cloned and how to price AI-generated content that replaces human labor.