newsfilter.io
Conference Presentation

How to Scale AI Application Inference 100x ft. Fireworks’ Lin Qiao

  • Future inference alignment will occur outside application space, requiring the field to ensure the process is easy, accessible, and high quality.
  • Scaling laws in inference will evolve to optimize simultaneously across three dimensions: quality, speed, and user concurrency (cost), rather than focusing on a single metric.
  • The industry approach will shift to heavy customization of scaling inference as a multi-dimensional optimization problem, targeting specific applications rather than generic models.
  • To resolve a combinatorial explosion of over 100,000 combinations, the focus will move to co-optimizing post-training and inference to drive model innovation and acceleration.
  • The ultimate objective is to reduce current extremely high inference costs by 10 to 100 times, enabling applications with product-market fit to scale sustainably.
  • Solving the three-dimensional optimization requires predicting 10 tokens at a time rather than one, while aligning numerics, positions, hardware selection, model sharding, and tuning mechanisms to specific application data distributions.
  • Significant R&D investment is directed toward solving this complex combinatorial problem, with a platform that abstracts hardware management complexity across multiple vendors and SKUs to ensure high quality and reliability.
  • The platform enables users to tune for speed or quality using distinct mechanisms to select optimized kernels and apply specific tuning strategies based on application needs.
  • Expectations include the incorporation of production data for reinforcement tuning to align models with specific application requirements.
  • A developer-facing self-serve platform has been available for the past year and a half, with an expectation for more applications to onboard and conduct experimentation using this alignment lens.
  • The platform aims to guide companies to the optimal balance, or "sweet spot," across quality, cost, and speed as they scale.
  • Further discussion is anticipated regarding next-generation inference problems based on the specific applications of attendees.