Fireside Chat, Interview
The Story Behind ElevenLabs | Interview with CEO Mati Staniszewski
- Voice is expected to become a fundamental human-computer interface, with screen-first activities on laptops and phones shifting to the background to enhance user presence.
- Future applications include immersive educational tools acting as expert tutors and cross-cultural experiences that enable users to understand both linguistic content and delivery nuances.
- The company observed user base growth from a few thousand to a few hundred thousand users shortly after launching in early January, a rate described as higher than initial expectations.
- Eleven Labs currently employs over 300 people across 11 cities and is doubling its team size every six months while hiring top-tier voice researchers globally.
- Product and research teams will immediately iterate together, layering research on specific value propositions to solve existing problems where companies currently separate these functions.
- A long-term objective involves developing a single model capable of generating any audio type, including converting voice into music, singing, or sound effects.
- The organization aims to become the first to pass the vocal Turing test with an empathetic, smart AI that facilitates human-like interaction.
- Most future human-machine communication is predicted to occur via audio due to its speed and information density compared to text.
- General audio generation models trained on raw data are expected to make machines significantly more capable across various raw data domains than text-trained models.
- The company strives to make its technology available globally across all languages and geographies to serve as the primary voice of technological change.
- Risks and validation plans rely on specific use cases that will resonate with the market following signals received after the January launch.