newsfilter.io
Interview, Fireside Chat

AssemblyAI Now Handles 4x YouTube's Daily Volume

  • Weekly API conversations at Assembly are projected to grow, with peak weeks expected to exceed 120 million voice conversations and 2 million hours of voice volume, a figure noted as more than four times the daily volume of YouTube as of the previous December.
  • The Total Addressable Market (TAM) is predicted to expand by 100x as potential builders shift from engineering teams to "anyone" utilizing coding agents, with global enterprises increasingly building custom software internally where teams as small as two people can now deploy solutions on the platform.
  • Voice AI is forecast to become a core component of software and hardware over the next 10 years, following a trajectory similar to self-driving cars where adoption continuously crosses thresholds as technology improves, while the next 12 to 18 months will see a significant increase in consumer applications and hardware integrating voice as a core dimension, including toys, games, and consumer electronics featuring on-device models.
  • Over the next couple of years, voice is expected to be integrated into consumer electronics to enable more passive computing experiences that free users from screen dependency, with a prediction that within five years, children will expect to talk to physical objects and devices as a standard norm comparable to current touchscreen usage.
  • The industry is expected to focus on the "get this stuff to work" phase for the next 18 months to resolve disambiguation problems in voice agents, particularly regarding multiple speakers and background noise, as new context-aware models capable of isolating specific speakers will significantly improve accuracy and enable wider deployment.
  • Future model evolution will address the need for a single model to handle 20 different languages with real-time switching and specific alignment needs for different cultures, while the industry must also determine the UX equilibrium between human mimicry and AI disclosure to avoid user confusion.
  • Voice is currently considered a reliable form of data capture for use cases such as medical chart updates, field service coaching, and automated note-taking, driven by improvements in the broader AI infrastructure ecosystem including reasoning models, vector databases, and the rise of coding agents.
  • Customers previously using open-source models are expected to return to the platform due to the maintenance difficulties and technological lag of open-source voice solutions, while big consumer electronics companies are actively implementing the technology for real-time command understanding.
  • The company anticipates reaching a size of 80 people soon, maintaining a small team structure to leverage AI agents for scale, with model updates expected to be released every couple of weeks to ensure constant improvement rather than infrequent annual releases.
  • Small businesses are expected to utilize coding agents such as Lovable, Cloud Code, Cursor, and Replit to build on the API infrastructure, a trend noted as unanticipated two years ago.
  • Long-term goals include the eventual translation of animal communication, though the speaker notes this interface may require sub-vocal or mind-reading technology rather than current voice capabilities.