newsfilter.io

Stanford CS153 Frontier Systems | Mati Staniszewski from ElevenLabs on The Future of Voice Systems

  • Community-adopted technology is expected to diffuse to the broader market within 6, 12, or 18 months, revealing new use cases.
  • Expressivity and emotionality in voice agents are projected to be largely resolved later this year, enabling emotion detection and tonal matching.
  • A "middle-to-middle" workflow where AI assists iterative refinement is anticipated to mitigate concerns regarding "AI slop," while non-execution interactions may utilize fused approaches before switching to cascaded architectures for authenticated transactions.
  • The cascaded model architecture is expected to remain the preferred enterprise standard for reliability over the next few years, with internal exploration of both cascaded and fused methods continuing to determine the optimal path for different interaction types.
  • Fused model approaches are predicted to suit speed-prioritized workflows and may become viable for non-execution interactions, though training remains complex due to token fusion difficulties and reliance on open-source intelligence.
  • By 2026, models are expected to evolve into fused or cascaded systems designed to further reduce latency and enhance real-time response capabilities.
  • Foundational research into audio, including conversational models and multimodal fusion, is planned for the next five years.
  • The distinction between platforms and applications is expected to blur over time, allowing easier application creation on the platform.
  • Within a five-year horizon, the market is forecast to stabilize with three to five platforms facilitating conversational business-to-audience setups.
  • Studio adoption of AI voiceovers is expected to increase significantly in the last six months due to breakthroughs in directorial control over narrative delivery.
  • High-end Hollywood studios are predicted to potentially replace end-to-end workflows with AI voiceovers once quality thresholds are met and economic models for IP and usage rights are resolved.
  • The primary bottlenecks for the wider audio space beyond 2025 are identified as finding optimal compute amounts and solving challenges related to personalizing interactions to individual user preferences.
  • Emotionality and controllability will be baked into model training steps prior to pipeline combination to ensure performance reliability in cascaded architectures.
  • The "one person frontier team" model is expected to become more feasible as tools enable individuals to build state-of-the-art systems previously requiring large teams.
  • Revenue growth is expected to remain highly predictable on the enterprise side (constituting over 50% of revenue), whereas the PLG and self-serve side will remain less predictable due to reliance on continuous model innovation.
  • Team expansion across global bases in London, New York, Warsaw, and San Francisco will continue, maintaining small teams of under 10 people to preserve speed and ownership.
  • Support for Western allies and government initiatives in Ukraine will continue in alignment with legal guidelines, alongside efforts to make voice synthesis accessible to individuals affected by conditions like ALS or throat cancer.
  • Non-open source video models are anticipated to continue trending as labs catch up to the frontier, contrasting with Eleven Labs' approach of maintaining active open ecosystems.
  • Model quality will remain the priority for on-device or on-prem deployment, which is expected to be limited to text-to-speech narration lacking the full interactive and emotional capabilities of cloud versions, with a persistent quality gap regarding transcription and emotional transfer for the foreseeable future.
  • Equal capacity and model access will be maintained for solo creators compared to large companies, subject to minor concurrency limits and compliance adjustments.
  • Current concerns about non-open source video models like Seed Dance are expected to persist as labs close the gap to the frontier.