newsfilter.io
Fireside Chat, Interview

From Cloud to Edge: AI Gets Personal

  • Smaller generative AI models are expected to gain popularity on device within the next year, with 2 billion to 8 billion parameter models deemed sufficient for robust text, image, or audio generation.
  • Generative models capable of producing image, voice, and video content are predicted to become prevalent on devices and within applications over the next two years, similar to traditional machine learning models.
  • Real-time voice agent inference workloads are projected to begin running locally within the next 12 to 18 months.
  • Smartphone compute power is currently estimated to match that of computers from 10 to 20 years ago, while technology for AR experiences utilizing cameras as projectors is already available.
  • Diffusion models are described as intrinsically smaller than large text models, and distillation technologies allow for the maintenance of capabilities in smaller parameter sizes.
  • Inference pricing for larger models has been dropping significantly, and the hardware development sector is anticipated to see increased interest and enthusiasm regarding chips and tools.
  • The overall trend is expected to impact the entire supply chain over the long run as foundation model technology matures and infrastructure prepares for new consumer experiences.
  • On-device models are forecasted to play a major role in interactions with the 3D and physical world, although it remains uncertain whether they will substantially reduce infrastructure costs for certain applications.
  • Shifting to on-device deployment requires applications and hardware to update simultaneously, contrasting with the continuous launch cycles possible for cloud-based models.
  • Architecting the tool chain is expected to alter economics by improving developer efficiency and iteration speed despite the logistical complexities of on-device updates.
From Cloud to Edge: AI Gets Personal — Outlook