Fireside Chat, Interview
ElevenLabs' Mati Staniszewski: How Voice Becomes the Interface for AI
- Future audio domain capabilities will prioritize matching human emotional state and intonation, with audio general intelligence enabling continuous narration, pausing, and singing becoming possible very soon.
- Voice is predicted to become the primary interface for interacting with ubiquitous robots and technology, particularly in sectors undergoing transformation by government support, citizen education, healthcare, and security infrastructure management.
- Agent capabilities will evolve toward revenue-generating sales applications (inbound and outbound) and interactive education features available 24/7, while moving beyond simple phone tree replacements.
- Technical resource requirements in non-technical teams like legal and sales will drive process automation and security management, with agents increasingly capable of interrupting human negotiations.
- Agent-to-agent communication is projected to occur at 100%, potentially shifting to efficient non-spoken information transfer, while human-to-human interaction value is expected to rise as voice technology becomes more common.
- Authentication protocols will shift to assume all non-verified entities are fake, requiring encoding and decoding for authenticated humans, while models will increasingly rely on audio as a minor stack component focused on workflow and ecosystem building.
- Investment in labeling audio data regarding emotional delivery and voice descriptions is planned for the next 6 to 12 months to deliver value within 10 to 24 months.
- An ecosystem aiming to provide access to over 20,000 contributed voices across diverse language styles and cultures is being built, though immediate challenges remain for voice agents in true emotional interactions and top-chart music production, with music capabilities expected to evolve over the next one to two years.
- The speaker intends to continue methodically opening up untapped niches in the audio domain.