newsfilter.io
Interview, Fireside Chat

A fireside Chat with Mati Staniszewski, CEO & Co-Founder of Eleven Labs ft. Balderton Capital

Company Origins and Strategic Vision

  • Eleven Labs was founded 15 years ago by co-founders Mati Kovacs and Peter in response to the poor quality of Polish film dubbing, where a single narrator often voiced all characters.
  • The founders pivoted from early hackathon projects (crypto risk analysis, speaking coaches) to Eleven Labs after recognizing voice as a critical, underserved interface in early 2021.
  • The company operates on a dual strategy of conducting frontier AI research while simultaneously building consumer and enterprise applications.
  • Mati Kovacs identified three core drivers for the company's expansion: converting text-based content to audio, enabling global content dubbing, and shifting technology interfaces to be voice-first.
  • Eleven Labs has established global offices in the US, Europe, India, and Japan to support its go-to-market strategy with localized language and cultural understanding.

State of the Art and Performance Metrics

  • The company's internal goal for the current year is to solve the "voice Turing test" for most conversation settings, enabling natural interruptions and human-like latency.
  • In information-heavy use cases (e.g., customer support widgets), voice agents already perform as well as or better than human operators.
  • Eleven Labs reports that 80% of its voice agents in healthcare scheduling (partnering with Elise AI and Hippocratic) are currently in production, automating appointment handling and patient check-ins.
  • The platform currently supports over 5,000 user-created voices, with the company having paid over $5 million in compensation to its creator community.
  • A key differentiator is the platform's ability to distinguish between dialects, such as Spanish-Mexican, Spanish-European, and Catalan, which significantly impacts completion rates in call centers.
  • Eleven Labs recently partnered with Epic Games to enable millions of Fortnite players to interact with a live, personalized Darth Vader companion in-game.

Safety, Provenance, and Future Interface Standards

  • To combat misuse, Eleven Labs employs three safety pillars: content provenance (tracing audio generation), fraud detection/moderation, and open-source classifier development for AI detection.
  • Kovacs proposes a future three-tier trust system for voice interactions:
    • Tier 1: Validated real human speech encoded/decoded on-device (e.g., via MAPI).
    • Tier 2: Approved, watermarked AI speech with explicit consent (e.g., celebrity likenesses like James Earl Jones for Star Wars).
    • Tier 3: Default unverified AI content requiring user skepticism.
  • The company anticipates a societal shift where users increasingly expect and adapt to AI voices as the primary interface for daily interactions like restaurant bookings and customer support.

Model Architecture and Technical Trajectory

  • Eleven Labs is currently splitting development between cascaded models (STT -> LLM -> TTS) and end-to-end "Duplex" speech-to-speech models.
  • The company experienced a strategic pivot in the last 18 months, initially over-investing in Duplex models before realizing cascaded models currently offer superior reliability for enterprise use cases.
  • The strategic forecast suggests cascaded models will dominate the next 12 months for high-reliability enterprise tasks, while Duplex models will lead in latency-sensitive personal assistant applications.
  • The company aims to deploy a reliable, true Duplex model within the current year to combine the latency benefits of speech-to-speech with the reliability of cascaded systems.

Market Applications and Forward-Looking Goals

  • Eleven Labs views "voice agents" as the primary vehicle for elevating brand interaction and accessibility in customer support and service discovery.
  • The company identifies AI-powered personal education (tutors) as its most socially impactful, though commercially complex, near-term use case, aiming to replicate the benefits of one-on-one instruction.
  • Investors and the company see immediate value in replacing frustrating voice interactions, such as traditional call centers, with snappy, AI-optimized alternatives.
  • The platform continues to expand its marketplace, allowing individuals to license their voices or create unique synthetic voices for compensation.