newsfilter.io
Conference Presentation, Product Demonstration

Lin Qiao, Co-Founder & CEO, Fireworks AI: Open Models: The future of AI agent development

  • 2025 is projected as the year of AI agents, with widespread testing and planning across verticals including retail, insurance, finance, education, medical, manufacturing, and security.
  • A predicted industry disruption stems from data misalignment between frontier model developers and application companies, expected to make off-the-shelf models less accurate, slower, and more costly, driving adoption of the "data flying wheel" pattern where product analytics train models to match data distribution.
  • Fireworks aims to facilitate this shift by offering tools for agentic and application developers to easily create data flying wheels, leveraging a platform that grew from a small number of models to 200 over the past year and a half, with recent releases enabling selection from thousands of models for agentic use cases without limitation.
  • Agentic development is anticipated to scale to millions of developers or billions of consumers in a very short timeframe, driven by the need for fast iteration loops; the newly General Available experimentation platform addresses this by allowing access to over 1,000 models with fast, powerful fine-tuning and zero setup.
  • The Multilora feature is designed to compress GPU footprints by 100 times, enabling developers to cycle through 100 times more experimentation and eliminating wait times for GPU resources.
  • Fireworks has expanded its model catalog to include audio in and out capabilities, alongside state-of-the-art models such as LAMA, Big Seq, DeepSeq, QAN, and five others, with new capabilities for long contexts exceeding 100,000 tokens and reinforcement fine-tuning to explore model variation.
  • Combined supervised and reinforcement fine-tuning is recommended for optimal results, with examples showing accuracy lifts of 18 points in domain reasoning and a specific code rewrite model achieving 15 points higher accuracy, 50 times fewer errors, and eight times faster performance than frontier lab models.
  • The platform's virtual cloud abstracts eight different cloud providers to ensure enterprise-grade reliability with 24/7 high SLA, supporting BYOC storage for security, while currently operating at 100,000 requests per second and processing over 5 trillion tokens daily.
  • A proprietary 3D optimizer handles latency, cost, and quality across 100,000 possible hardware combinations, delivering 8 to 15 times acceleration for co-react use cases and 4 to 8 times acceleration for deep research use cases.
  • Rapid scaling is evidenced by an enterprise customer growing from one to 100 stores in three days and aiming for 1,000 stores in three months, a digital native scaling from 1 million to 30 million users in one month, and a software productivity firm scaling from 10 million to 100 million users within a month.
  • Fireworks reports onboarding over 10,000 companies, representing 18 times growth within a year, and is offering 5,000 credits to attendees to build on the platform.