newsfilter.io
Conference Presentation, Lecture

Post-Training Is How You Keep Your Taste | Fireworks CEO Lin Qiao

  • Companies are expected to shift from relying on off-the-shelf APIs to developing deeper, specialized models via post-training to encode unique judgment and taste, with a potential future landscape of millions of models optimized for individual use cases.
  • Post-training is positioned as a strategic vehicle to reduce operational costs by five to ten times, enabling significantly higher traffic volumes and avoiding "scaling into bankruptcy" while improving unit economics.
  • Organizations are advised to begin experimenting with post-training early but may specifically prioritize it after achieving product-market fit to leverage meaningful, high-quality data, with the process potentially requiring a one to two-year planning horizon before usage spikes occur.
  • The progression of AI capabilities is predicted to move from prompts to RAG for dynamic facts, supervised fine-tuning for output corrections, preference tuning (DPO) for product-specific taste, and reinforcement learning for specific problems or weak areas.
  • Technical risks include reward hacking due to misaligned RL environments, quality drops from training-serving stack misalignment, and precision loss if libraries or numerics differ, necessitating the conversion of judgment into systemic evaluation frameworks.
  • Successful implementation requires convergence between product and ML teams to manage data quality, with the potential to block AI feature launches by CFOs due to cost burdens unless remediated.
  • Companies may pursue varied technical architectures, ranging from building bespoke customized harnesses with low-level API control to using training SDKs with deployed researchers, depending on their need for granular control versus engineering support.
  • Specialized models are anticipated to be built for domain-specific sectors such as legal, finance, healthcare, and marketing, where tuned models may outperform state-of-the-art closed models if frontier labs lack specific industry data.
  • Financial and operational realities involve an initial intercept of investment and ongoing maintenance costs, with the need to blend unique reward criteria into a "Secret Sauce" and the risk of burnout if ideal results are not reached despite time and money expenditures.
  • Future competitive dynamics may see companies announcing new models every few months to compete at frontier quality levels while maintaining lower costs, potentially leading to a market of bespoke models that cannot be easily cloned.