Interview, Fireside Chat
The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao
Investment Thesis & Market Outlook
- Lin Kuao asserts that the current year (following "The Year of Coding") is "The Year of Cowork," marking a shift toward specialized intelligence applications.
- He predicts a 10x reduction in token costs within the next three years, which will subsequently drive 100x growth in usage volume.
- Kuao explicitly states that Fireworks will not move into the application layer, maintaining a focus on the infrastructure and specialized intelligence platform.
- The company is currently open to exploring data center capabilities in the future, contingent on timing and scale, though they prioritize agility over vertical integration at this stage.
- Kuao forecasts that the future of intelligence will not be dominated by a single "AGI" model but rather millions of specialized models, one for each specific use case.
- He anticipates that by year-end, Fireworks will double its Annual Recurring Revenue (ARR) to at least $1.6 billion from a current base of $800 million.
- The company currently processes over 40 trillion tokens per day, with the majority derived from customized models rather than off-the-shelf variants.
Strategic Philosophy: Specialization vs. Generalization
- Kuao argues against a "one-company-owns-intelligence" model, viewing it as a threat to human creativity and diverse value systems.
- He believes that the majority of valuable data (private enterprise data) is not used for training general models and that the future lies in "activating" this proprietary data for specialized intelligence.
- Unlike companies like Anthropic that pursue a single general-purpose AGI, Fireworks focuses on specialized intelligence to align with the unique "taste," policies, and judgment of individual enterprises.
- Kuao posits that open-source models are becoming a critical alternative to frontier models, offering full control over weights, zero acquisition cost, and the ability to steer models with small amounts of unique data.
- He notes that while frontier models provide essential "power lines" for the economy, they cannot replace the need for specialized, customized intelligence for specific business workflows.
Operational Dynamics & Growth
- Fireworks grew from a handshake investment to $800M ARR in a span of roughly two years, a growth velocity Kuao attributes to the unique disruption of the AI stack.
- The company recently hired George Huang (former President of Salesforce) to scale its GTM team, citing the need to match a super-linear demand curve with increasing team productivity.
- Kuao admits to initially underestimating the speed of growth and delaying hiring, but has since shifted strategy to aggressively use AI tools and hire for "extreme ownership."
- The company's gross margin structure is currently higher than traditional SaaS (often 30-40%) due to the hyper-growth phase, which prioritizes speed and innovation over margin optimization.
- Fireworks targets high-velocity growth by hiring "high immigrants" (a specific demographic focus) who demonstrate rigorous data usage, accountability, and an "unwavering sense of ownership."
Technology Stack & Infrastructure
- Fireworks distinguishes itself through "Zero KLD" (Kullback-Leibler Divergence) training, ensuring numerical equivalence between training and inference to prevent quality loss during deployment.
- The company utilizes a distributed system design across multiple data center regions to manage massive RL (Reinforcement Learning) jobs without relying on a single, expensive hyperscaler cluster.
- Kuao emphasizes that hardware depreciation cycles (now accelerating to 9 SKUs per year per vendor) have outpaced model iteration, complicating the "build vs. buy" decision for custom chips.
- He identifies the lack of system-level design for models exceeding 10 trillion parameters as a current critical bottleneck, requiring co-design from the model layer down to the chip layer.
- The company views the routing layer (automated selection of models based on task complexity) as a future area of high innovation and potential automation.
- Kuao believes that while major tech companies (NVIDIA, Meta, DeepSeek) are building chips, it is premature for startups to do so until workloads stabilize, as hardware is difficult to modify once taped out.
Competitive Landscape & Risks
- Kuao views the legal AI sector (e.g., Harvey vs. Legora) as a prime example of companies needing to own their intelligence to handle conservative, error-intolerant workflows and proprietary data.
- He suggests that while open-source models (including Chinese models) are rising, US open ecosystems will eventually rebuild to ensure supply chain security and control.
- Concerns regarding national security and "sovereign models" are validated; Kuao predicts nations and blocks will eventually own their own infrastructure to avoid dependency risks similar to power grid outages.
- He predicts that 80-90% of enterprise workflows will eventually be handled by tuned open models, reducing the reliance on expensive frontier model APIs for general tasks.
- The industry is transitioning from "token maxing" (optimizing token efficiency) to "ROI maxing" (optimizing return on investment), necessitating better tools for monitoring business value attribution.
Leadership & Culture
- Kuao highlights that leadership in AI requires deep, ground-level judgment rather than privilege, citing Jensen Huang's direct engagement with details as a model for maintaining velocity.
- He admits that the company initially under-prioritized marketing and customer education, focusing instead on product development, but now recognizes the need to drive clarity and adoption.
- The "build vs. buy" decision for hardware is framed as a timing issue; companies should wait until workloads are stable and the cost savings (e.g., 5x reduction) justify the risk of building custom silicon.
- Kuao predicts that by the end of the next year, every company will view owning their own intelligence as a non-negotiable necessity, similar to building their own software stack in the SaaS era.