Conference Presentation, Product Demonstration
Lin Qiao, Co-Founder & CEO, Fireworks AI: Open Models: The future of AI agent development
Market Positioning & Vision
- CEO Lin defines 2025 as "the year of AI agents," predicting a surge in production deployments across retail, insurance, finance, education, medical, manufacturing, and security.
- Fireworks powers production agents for major enterprise clients, including:
- Cursor: Utilizing Fireworks for its flagship coding agent to accelerate software development.
- Uber: Deploying agents for driver onboarding, personalized customer support, and "vibe-SQLing" (text-to-SQL analytics).
- UiPath: Executing a strategic pivot to launch an "agentic builder" product combining AI robots with human workflows.
- Strategic Prediction: Companies that successfully align their proprietary product data distributions with their underlying model distributions will disrupt their respective industries.
Core Technical Philosophy: The Data Flying Wheel
- Problem: Off-the-shelf models (closed or open) suffer from data misalignment, resulting in lower accuracy, slower inference, and higher costs compared to custom-tuned solutions.
- Solution: A "data flying wheel" mechanism where production analytics logs feed back into model training to align the model's data distribution with the product's specific domain.
- Value Proposition: Continuous alignment improves model quality, which in turn improves product quality, creating a compounding competitive advantage.
Platform Capabilities & Scale
- Model Catalog: Expanded from a limited set to over 1,000 models, including LLMs, image generation, embeddings, vision, and audio (in/out).
- Growth Metrics: Onboarded >10,000 companies, representing 18x growth in one year.
- Infrastructure Scale:
- Processes 5 trillion tokens per day (3x Microsoft Azure's reported volume).
- Handles ~100,000 requests per second (traffic volume comparable to Google Search).
- Virtual Cloud: Abstracts hardware and cloud provider complexity, running on eight different providers with disaggregated inference and reinforcement tuning engines.
- 3D Optimizer: An automated system searching through 100,000+ combinations to optimize simultaneously for:
- Quality.
- Latency (e.g., 500ms targets).
- Cost (e.g., 10x reduction targets).
- Performance Gains: Achieved 8-15x acceleration for coding cases and 4-8x for deep research cases compared to baselines.
Development Workflow & Tooling
- Multilora: A feature compressing GPU footprint by 100x to enable rapid experimentation, allowing developers to fine-tune and iterate 100x faster without hardware bottlenecks.
- Customization Strategies:
- Supervised Fine-Tuning (SFT): Supports state-of-the-art open models (Llama, BigSeq, DeepSeq, Qwen) with long context windows (>100k tokens) and quantization for speed.
- Reinforcement Fine-Tuning (RFT): Introduces reward-based training (positive/negative penalties) to explore model search space; accessible via a Web IDE and Reward Kit SDK.
- Synergistic Approach: Recommends starting with SFT to establish a baseline, followed by RFT to refine specific domain behaviors.
Quantitative Performance Case Studies
- Code Rewriting: Reinforcement fine-tuning an open model yielded a 15-point accuracy gain over the baseline, resulting in 50x fewer errors and 8x faster inference than frontier models.
- Reasoning Tasks: Combining SFT and RFT lifted accuracy by 18 points.
- Function Calling: Open models were tuned to match the accuracy of frontier labs while maintaining smaller model sizes.
Customer Scaling Success Stories
- Retail Chain: Scaled deployment from 1 store to 100 stores in three days; currently scaling to 1,000 stores within three months.
- Digital Native: Expanded an AI-native feature from 1 million to 30 million users in one month.
- Software Productivity: Scaled from 10 million to 100 million users within one month.
Enterprise & Security Features
- Designed with enterprise-grade security, privacy, and compliance readiness.
- Supports Bring Your Own Cloud (BYOC) and storage buckets for customers requiring strict data sovereignty.
- Offers a Build SDK for programmatic scaling of GPU fleets and system experimentation.
Call to Action
- Fireworks is offering 5,000 credits to attendees to facilitate immediate experimentation and building.
- Developers are encouraged to treat the model as a core IP asset rather than a utility, leveraging Fireworks tools to build, customize, and scale rapidly.