Interview
Inference's Scaling Era | Tuhin Srivastava, Baseten & Corinne Riley, Greylock | RAISE Summit 2026
Market Projections
- Thuy Hinh, CEO of Base10, forecasts inference will become one of the world's largest markets by 2030.
- Based on a projected $5 trillion global AI spend in 2030, Hinh estimates approximately $2.5 trillion will be allocated specifically to inference costs.
- This projection assumes the application layer achieves roughly 50% gross margins, with inference serving as the cost of goods sold (COGS) for AI value generation.
Base10 Infrastructure Strategy
- Base10 positions itself as an "inference cloud" comprising compute, software, and ancillary services including RL post-training, evals, sandboxes, and routing.
- The company operates clusters across 20 different clouds in 90 distinct regions to ensure scalability and latency optimization.
- Base10 focuses on reducing the burden on customer infrastructure teams by providing full-stack software for inference and post-training rather than just bare metal hardware.
Enterprise Financial Impact & Adoption Trends
- Introducing AI to enterprise applications has been observed to slash gross margins by 20 to 30 percentage points for companies with $500 million to $1 billion in revenue.
- Current market data indicates 3% to 5% of total token spend is currently on open-source and custom models, while sophisticated customers already allocate 40% to 50% of their spend and volume to these models.
- Decagon publicly reported that 90% of its model calls now utilize open-source models.
- Hinh predicts the average split between closed and open-source token volume will reach 40-50% within the next year, driven by the ability to run more inference at lower costs.
Open-Source Landscape & Technology
- Open-source models now lag behind frontier closed-source models by only approximately three months, a gap Hinh believes will continue to narrow.
- The gap is narrowing due to increased capital investment in talent and compute, exemplified by NVIDIA's investment in the open-source Nemotron model.
- Post-training on open-source data allows enterprises to achieve 95% to 100% of frontier model intelligence at roughly 70% lower costs.
- Bridgewater is cited as a recent example of a firm outperforming closed labs on frontier tasks through post-training techniques.
- Specific high-volume use cases for open-source models include code generation, which Hinh notes is currently 77% powered by open-source models.
Forward-Looking Statements & Regulatory Outlook
- Hinh anticipates that a Fortune 500 company will reach 90% open-source model adoption within the next 12 to 18 months.
- Future growth drivers are identified as "infinite generation" scenarios, specifically within biology and pharmaceutical sectors for compound generation.
- Despite regulatory discussions, Hink expresses confidence that open-source will continue to grow, citing the TCP/IP internet protocol as a blueprint for a competitive ecosystem.
- Enterprises are driven to adopt open-source to retain ownership of their data, workflows, and customer signals, preventing dependency on external labs that could "yank away" access.
Customer Decision Factors
- Selection between inference providers is described as a "multifaceted scorecard" rather than a single KPI, weighing speed, latency, reliability, and post-training capabilities.
- Customers are increasingly distinguishing providers based on the ability to offer a "verticalized cloud" with integrated tools for search, tool calling, and evaluation.