Fireside Chat, Interview
The Story Behind ElevenLabs | Interview with CEO Mati Staniszewski
Market Evolution and Future Interfaces
- Human voice synthesis efforts date back to the 1700s, with digital synthesizers emerging in the early 1900s, yet few technologies previously crossed the threshold to sound convincingly human or evoke emotion.
- Voice is projected to become the next fundamental computer interface, surpassing screens, touch, and keyboards as the primary mode of interaction.
- This shift is expected to move computing from screen-first to background processing, allowing users to remain more present (e.g., students utilizing smart physics or history assistants via headphones).
- Voice technology aims to eliminate language and cultural barriers, enabling full immersion in foreign cultures through accurate translation of both semantic content and intonation.
- Most future human-machine communication may occur via audio, as it is faster to transmit and more information-rich than text, capturing nuances that text-based LLMs miss.
- Audio is the only AI modality capable of eliciting genuine emotional responses (e.g., ASMR, cinematic voices) that text-based interactions cannot replicate.
Product Philosophy and Technical Capabilities
- Eleven Labs' guiding principle is the integration of proprietary research with direct product application, allowing real-time feedback loops where users define research needs and researchers iterate models on actual product data.
- The company aims to be the first to cross the "vocal Turing test," creating AI that is human-like, conversational, and empathetic.
- Current specialized models cover audio, sound effects, and music, with a strategic goal to develop a single generative model capable of synthesizing any audio type (e.g., converting voice to music or sound effects).
- Training a general audio generation model on raw audio data (rather than text tokens) is seen as the pathway to making AI smarter across all raw data domains.
- Recent product milestones include the introduction of Voice Design V3, Eleven Labs Image and Video, and Studio 3.0.
Company Origins, Growth, and Team Dynamics
- The concept for Eleven Labs originated in Poland due to the frustration with foreign media where all characters were voiced by a single narrator, stripping away emotion and intonation.
- Co-founders Jacek (Marty) and Piotr, previously at Palantir and Google respectively, began iterating on the product in 2021 on weekends before launching in early January with a waiting list of a few thousand users.
- User adoption rapidly scaled from a few thousand to hundreds of thousands, exceeding initial expectations by an order of magnitude.
- The company has grown from a two-person pre-seed team to over 300 employees across 11 cities, doubling in size every six months while maintaining a remote-first structure.
- Hiring strategy focuses on individuals with "proof of excellence" outside traditional backgrounds (e.g., astrophysics, open-source projects, high-level gaming rankings) rather than conventional credentials.
- The remote-first culture operates with high autonomy, small teams, and zero bureaucracy to preserve a global talent pool in a niche field where only ~50 to 100 top-tier voice researchers exist globally.
Organizational Structure and Culture
- The company has removed all job titles to filter for low-ego candidates and to eliminate implicit bias in asking for help or proposing ideas, ensuring a flat hierarchy.
- Operational scaling is driven by "mini-founders" who take ownership, supported by a culture of high trust and low bureaucracy.
- Co-founders Matti and Piotr function as a complementary pair (described as Yin/Yang), with Piotr focusing on deep technical research and Matti handling operations and vision; investors describe Piotr as technically genius and Matti as the "good cop."
- Cultural fit screening is rigorous to ensure the team can scale quickly while maintaining the unique, high-trust, family-like environment established by the founders.
- The founders' personal motivation is driven by the ability to define the future voice interface alongside a team of close friends and trusted colleagues.