Fireside Chat, Interview
Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Applications, Applied AI
Base10 & The Future of AI Inference: Key Insights
- Founder Background: Toohin (Toohan) is the CEO and co-founder of Base10, originally from Sydney, Australia, with a career starting in finance (Macquarie Bank) before pivoting to machine learning engineering in 2011–2012; he has been building in tech for over a decade, including early-stage ventures since 2015.
- Company Mission: Base10 provides production inference infrastructure designed to power custom, post-trained AI models, aiming to facilitate the shift from "95% spend on frontier models" to "5% on custom/open-source models" to improve margins and defensibility.
- Scale & Growth: The company powers some of the fastest-growing AI companies, handling approximately 30 trillion tokens per day, a volume that currently exceeds the combined daily token usage of OpenAI's API and Google's Gemini API.
- Strategic Thesis: The core belief is that inference demand will grow 1,000x to 1,000,000,000x (a "billion X" increase) due to the rise of agentic applications and larger models, making compute scarcity a permanent, non-resetting condition.
- Customer Strategy (Why Base10 vs. Hyperscalers):
- Performance & Optimization: Unlike raw cloud providers (AWS, CoreWeave, etc.), Base10 provides a software stack that handles complex optimizations for latency and throughput across multiple models.
- Reliability & Multi-Cloud: Base10 abstracts away compute source, stitching together 18 different clouds and 87 clusters to create a fault-tolerant, resilient infrastructure.
- Defensibility: Customers use Base10 to post-train open-source models to own their proprietary data and workflows, preventing reliance on frontier labs that could eventually replicate their unique value propositions.
Business Model & Economics
- Pricing Evolution: Currently, Base10 primarily operates on a compute markup model (charging more for managed NVIDIA H100/B200 chips than raw cost), but is transitioning toward token-based pricing to align with customer app-layer economics.
- Cost Advantages:
- Post-training open-source models are 70% to 90% cheaper than running frontier closed-source models.
- Open-source models lag behind frontier models by approximately 90 days.
- Vertical integration (owning hardware) is projected to be 30% cheaper than renting compute at current market rates.
- Compute Scarcity Reality:
- The market for GPUs is described as a "drug market" with no efficient pricing mechanism; spot prices are volatile and non-transparent.
- Lease renewals have seen double-digit price increases (e.g., a cluster renewal cited at $263/hour rising to $510/hour, a ~90% increase).
- Lead times for GPU procurement are currently 12 to 15 months out (e.g., Q2 next year for current ordering cycles).
- Infrastructure Capex: To meet projected demand (estimated at 150,000 B200 equivalents in two years), Base10 projects a compute spend of approximately $7 billion, necessitating a shift from purely renting to owning hardware to guarantee supply.
Technology & Hardware Ecosystem
- Hardware Strategy:
- The current fleet runs predominantly on NVIDIA hardware (H100, B200) due to the entrenched CUDA ecosystem and optimized runtimes (e.g., TensorRT-LLM, SGLang).
- Base10 anticipates a shift toward heterogeneous architectures separating "pre-fill" and "decode" tasks across different specialized chips.
- The company is skeptical of competitors like Tranium or TPUs achieving scale comparable to NVIDIA's supply chain and ecosystem depth in the immediate term.
- Post-Training Workflows:
- Base10 offers a full-stack service where customers provide a utility function (what to optimize, e.g., transcription errors), a base model, and proprietary data.
- Base10 provides the scaffolding to post-train the model and deploy it for inference, allowing customers to own their "intelligence" stack without building the underlying ML ops from scratch.
- Future Hardware Vision: Toohin proposes modular data centers (standardized "compute containers") as a future industry standard to industrialize the build-out of AI infrastructure, creating an API for compute similar to the standardization of shipping containers.
Market Dynamics & Geopolitics
- Open Source vs. Closed Source:
- Base10 bets that open-source models will remain viable and necessary; however, the current best open-source models originate from China (e.g., Moonshot, Alibaba, MiniMax) due to US labs' strategic decision to insource research.
- There is a concern that without US-led open-source development, the West faces a national security risk where intelligence costs 70–90% cheaper in the East.
- Toohin argues that the "two-company" outcome (OpenAI/Anthropic dominance) is unsustainable for the global ecosystem and advocates for a "many models" world to preserve an application layer.
- Competitive Landscape:
- Base10 views hyperscalers (AWS, Azure, GCP) and dedicated AI clouds (CoreWeave, Nebius) as competitors who often fail to solve the "inference stack" complexity, leading customers to Base10 after initial failures on raw compute.
- Cloud providers are increasingly building their own inference layers; Base10 remains open to partnering with them while emphasizing the stickiness of their software layer.
- Risk Factors:
- The three primary risks identified are: (1) the failure of the open-source application layer to mature, (2) the inability to secure sufficient compute access, and (3) the consolidation of intelligence power into one or two entities.
- Base10 believes compute scarcity will never normalize because inference demand compounds continuously without a "reset period" (unlike human labor or traditional infrastructure).
Forward-Looking Statements & Advice
- Next Best Idea: If not building Base10, Toohin would focus on energy and power infrastructure, specifically the creation of modular, standardized data centers to solve the physical constraints of the AI boom.
- Investment Stance: He is bullish on computing markets but cautions that the current market structure is immature and inefficient compared to mature utilities like electricity.
- Advice to Students:
- Do not feel pressured to specialize narrowly; the field changes rapidly, and one can become an expert in six months.
- Focus on studying whatever is fun and engaging, as career paths are fluid.
- There is significant opportunity in the financing and project financing of massive data center builds.
- Market Outlook: Inference is becoming the only remaining market with significant room for growth if the AGI thesis holds, as pre-training becomes a monopoly and application layers face pressure from frontier labs.