newsfilter.io
Conference Presentation, Product Demonstration

Lin Qiao, Co-Founder & CEO, Fireworks AI: Open Models: The future of AI agent development

  • Market Positioning & Vision

    • CEO Lin defines 2025 as "the year of AI agents," predicting a surge in production deployments across retail, insurance, finance, education, medical, manufacturing, and security.
    • Fireworks powers production agents for major enterprise clients, including:
      • Cursor: Utilizing Fireworks for its flagship coding agent to accelerate software development.
      • Uber: Deploying agents for driver onboarding, personalized customer support, and "vibe-SQLing" (text-to-SQL analytics).
      • UiPath: Executing a strategic pivot to launch an "agentic builder" product combining AI robots with human workflows.
    • Strategic Prediction: Companies that successfully align their proprietary product data distributions with their underlying model distributions will disrupt their respective industries.
  • Core Technical Philosophy: The Data Flying Wheel

    • Problem: Off-the-shelf models (closed or open) suffer from data misalignment, resulting in lower accuracy, slower inference, and higher costs compared to custom-tuned solutions.
    • Solution: A "data flying wheel" mechanism where production analytics logs feed back into model training to align the model's data distribution with the product's specific domain.
    • Value Proposition: Continuous alignment improves model quality, which in turn improves product quality, creating a compounding competitive advantage.
  • Platform Capabilities & Scale

    • Model Catalog: Expanded from a limited set to over 1,000 models, including LLMs, image generation, embeddings, vision, and audio (in/out).
    • Growth Metrics: Onboarded >10,000 companies, representing 18x growth in one year.
    • Infrastructure Scale:
      • Processes 5 trillion tokens per day (3x Microsoft Azure's reported volume).
      • Handles ~100,000 requests per second (traffic volume comparable to Google Search).
    • Virtual Cloud: Abstracts hardware and cloud provider complexity, running on eight different providers with disaggregated inference and reinforcement tuning engines.
    • 3D Optimizer: An automated system searching through 100,000+ combinations to optimize simultaneously for:
      • Quality.
      • Latency (e.g., 500ms targets).
      • Cost (e.g., 10x reduction targets).
    • Performance Gains: Achieved 8-15x acceleration for coding cases and 4-8x for deep research cases compared to baselines.
  • Development Workflow & Tooling

    • Multilora: A feature compressing GPU footprint by 100x to enable rapid experimentation, allowing developers to fine-tune and iterate 100x faster without hardware bottlenecks.
    • Customization Strategies:
      • Supervised Fine-Tuning (SFT): Supports state-of-the-art open models (Llama, BigSeq, DeepSeq, Qwen) with long context windows (>100k tokens) and quantization for speed.
      • Reinforcement Fine-Tuning (RFT): Introduces reward-based training (positive/negative penalties) to explore model search space; accessible via a Web IDE and Reward Kit SDK.
    • Synergistic Approach: Recommends starting with SFT to establish a baseline, followed by RFT to refine specific domain behaviors.
  • Quantitative Performance Case Studies

    • Code Rewriting: Reinforcement fine-tuning an open model yielded a 15-point accuracy gain over the baseline, resulting in 50x fewer errors and 8x faster inference than frontier models.
    • Reasoning Tasks: Combining SFT and RFT lifted accuracy by 18 points.
    • Function Calling: Open models were tuned to match the accuracy of frontier labs while maintaining smaller model sizes.
  • Customer Scaling Success Stories

    • Retail Chain: Scaled deployment from 1 store to 100 stores in three days; currently scaling to 1,000 stores within three months.
    • Digital Native: Expanded an AI-native feature from 1 million to 30 million users in one month.
    • Software Productivity: Scaled from 10 million to 100 million users within one month.
  • Enterprise & Security Features

    • Designed with enterprise-grade security, privacy, and compliance readiness.
    • Supports Bring Your Own Cloud (BYOC) and storage buckets for customers requiring strict data sovereignty.
    • Offers a Build SDK for programmatic scaling of GPU fleets and system experimentation.
  • Call to Action

    • Fireworks is offering 5,000 credits to attendees to facilitate immediate experimentation and building.
    • Developers are encouraged to treat the model as a core IP asset rather than a utility, leveraging Fireworks tools to build, customize, and scale rapidly.