Interview, Fireside Chat
Inference 101: SambaNova CEO Rodrigo Liang
Capital & Valuation Milestones
- Sambanova just closed the first tranche of a $1 billion fundraise at an $11 billion valuation, led by General Atlantic with participation from Seligman Ventures, T. Rowe Price, and Capital Group.
- The company has now raised a cumulative total of $2.5 billion, distinguishing itself as one of few chip startups to achieve multi-billion dollar fundraising.
- Investors are citing "premium inference" capabilities and the ability to scale quickly as key drivers for the capital infusion.
Strategic Shift to Inference Scaling
- The company's focus has evolved from training efficiency (SN10/SN20 chips) to inference optimization, driven by the shift from model development to mass-scale deployment by companies like Anthropic, OpenAI, and Gemini.
- Inference scaling presents distinct challenges compared to training, specifically regarding power density, data center footprint, and latency for millions of daily users.
- Sambanova's SN40 rack outperforms 130–140 kilowatt NVIDIA GPU racks using only 10 kilowatts, enabling the deployment of trillion-parameter models in a single air-cooled rack rather than dozens.
- The technology allows for "premium inference" by running models at full original precision without quantization, maximizing accuracy for large models (1T to 10T parameters) while maintaining speed.
Infrastructure & Deployment Architecture
- Unlike training clusters requiring massive, synchronized networks where a single failure impacts the whole system, Sambanova's inference architecture scales out with a "minimum quantum" of a single rack.
- The company is promoting a heterogenous data center model where their 10kW racks can be deployed in existing standard 19-inch, air-cooled facilities, avoiding the 18-month timeline and massive CapEx of building gigawatt-scale, liquid-cooled data centers.
- Strategic partnerships, such as the NeoCloud "Vector Core Compute (VC2)" with Vista Equity and Cambium, allow Sambanova to ship infrastructure while partners handle data center construction and operations.
- Modular deployments are being tested in shipping containers for edge use cases, including remote energy sectors (oil rigs, mining) and military applications, where traditional power and space are unavailable.
Market Dynamics & Competitive Landscape
- The industry is entering a "land grab" phase where speed of scaling and user acquisition are the primary differentiators, as the market is expected to consolidate around 2–4 dominant infrastructure providers.
- Sambanova positions its chips as a co-opetition asset, allowing cloud providers to route high-margin inference traffic to their racks while keeping other racks for training or HPC, thereby improving provider margins.
- Customers are measuring success via "revenue per rack," calculated by tokens generated per second multiplied by the price per token, minus operational costs.
- The company predicts a bifurcation in the market where fast, accurate inference becomes the standard expectation, similar to the shift from 2G to 5G, eventually forcing consumers and enterprises to adopt high-speed tiers as costs decrease.
Sovereignty, Edge, and Future Trends
- Data sovereignty is driving demand for on-premise and national-level models, particularly in Europe, to prevent sensitive corporate and government data from being ingested into global public models.
- The "agentic" future of AI, where multiple agents orchestrate complex tasks (e.g., banking transfers, itinerary planning), necessitates ultra-low latency (sub-0.1 second response times per agent) which favors distributed, city-center data centers over centralized remote ones.
- The market is shifting from "AI for cost savings" to "AI for differentiation," where companies train proprietary models on private data to create unique services that competitors cannot replicate.
- Sambanova's upcoming Generation 5 chip (DSM50) is expected to further aggregate inference efficiency, supporting both hyperscale clusters and highly energy-efficient edge deployments.
Founder Perspective
- Rodrigo Yang, with 32 years in high-performance chip development, emphasizes that the current semiconductor interest is historically unprecedented, viewing chips as the central bottleneck for the AI transformation.
- He argues that success requires resilience and the ability to navigate cyclical industry phases (e.g., the return of on-premise infrastructure after decades of cloud centralization) to build enduring businesses.