newsfilter.io
Interview, Fireside Chat

Inference 101: SambaNova CEO Rodrigo Liang

  • Capital & Valuation Milestones

    • Sambanova just closed the first tranche of a $1 billion fundraise at an $11 billion valuation, led by General Atlantic with participation from Seligman Ventures, T. Rowe Price, and Capital Group.
    • The company has now raised a cumulative total of $2.5 billion, distinguishing itself as one of few chip startups to achieve multi-billion dollar fundraising.
    • Investors are citing "premium inference" capabilities and the ability to scale quickly as key drivers for the capital infusion.
  • Strategic Shift to Inference Scaling

    • The company's focus has evolved from training efficiency (SN10/SN20 chips) to inference optimization, driven by the shift from model development to mass-scale deployment by companies like Anthropic, OpenAI, and Gemini.
    • Inference scaling presents distinct challenges compared to training, specifically regarding power density, data center footprint, and latency for millions of daily users.
    • Sambanova's SN40 rack outperforms 130–140 kilowatt NVIDIA GPU racks using only 10 kilowatts, enabling the deployment of trillion-parameter models in a single air-cooled rack rather than dozens.
    • The technology allows for "premium inference" by running models at full original precision without quantization, maximizing accuracy for large models (1T to 10T parameters) while maintaining speed.
  • Infrastructure & Deployment Architecture

    • Unlike training clusters requiring massive, synchronized networks where a single failure impacts the whole system, Sambanova's inference architecture scales out with a "minimum quantum" of a single rack.
    • The company is promoting a heterogenous data center model where their 10kW racks can be deployed in existing standard 19-inch, air-cooled facilities, avoiding the 18-month timeline and massive CapEx of building gigawatt-scale, liquid-cooled data centers.
    • Strategic partnerships, such as the NeoCloud "Vector Core Compute (VC2)" with Vista Equity and Cambium, allow Sambanova to ship infrastructure while partners handle data center construction and operations.
    • Modular deployments are being tested in shipping containers for edge use cases, including remote energy sectors (oil rigs, mining) and military applications, where traditional power and space are unavailable.
  • Market Dynamics & Competitive Landscape

    • The industry is entering a "land grab" phase where speed of scaling and user acquisition are the primary differentiators, as the market is expected to consolidate around 2–4 dominant infrastructure providers.
    • Sambanova positions its chips as a co-opetition asset, allowing cloud providers to route high-margin inference traffic to their racks while keeping other racks for training or HPC, thereby improving provider margins.
    • Customers are measuring success via "revenue per rack," calculated by tokens generated per second multiplied by the price per token, minus operational costs.
    • The company predicts a bifurcation in the market where fast, accurate inference becomes the standard expectation, similar to the shift from 2G to 5G, eventually forcing consumers and enterprises to adopt high-speed tiers as costs decrease.
  • Sovereignty, Edge, and Future Trends

    • Data sovereignty is driving demand for on-premise and national-level models, particularly in Europe, to prevent sensitive corporate and government data from being ingested into global public models.
    • The "agentic" future of AI, where multiple agents orchestrate complex tasks (e.g., banking transfers, itinerary planning), necessitates ultra-low latency (sub-0.1 second response times per agent) which favors distributed, city-center data centers over centralized remote ones.
    • The market is shifting from "AI for cost savings" to "AI for differentiation," where companies train proprietary models on private data to create unique services that competitors cannot replicate.
    • Sambanova's upcoming Generation 5 chip (DSM50) is expected to further aggregate inference efficiency, supporting both hyperscale clusters and highly energy-efficient edge deployments.
  • Founder Perspective

    • Rodrigo Yang, with 32 years in high-performance chip development, emphasizes that the current semiconductor interest is historically unprecedented, viewing chips as the central bottleneck for the AI transformation.
    • He argues that success requires resilience and the ability to navigate cyclical industry phases (e.g., the return of on-premise infrastructure after decades of cloud centralization) to build enduring businesses.