newsfilter.io
Interview, Fireside Chat

Inference 101: SambaNova CEO Rodrigo Liang

  • A capital infusion from General Atlantic, Seligman Ventures, T. Rowe Price, and Capital Group is expected to drive the company toward scale and momentum.
  • Semiconductor interest is predicted to be at a 32-year high, shifting industry focus from model training to inference, with chip deployment numbers for inference anticipated to be orders of magnitude larger than for training.
  • The company plans to release its seventh chip, Generation 5, later this year, having previously taped out six chips over the last seven years, with the DSN 50 chip expected to enable both hyperscale clusters and energy-efficient edge deployments.
  • Inference demand is forecast to require efficiency improvements, with SambaNova predicting that a single 10 kilowatt air-cooled rack can run a trillion-parameter model, reducing the minimum cluster size from 10–20 racks to just one.
  • Deployment timelines for new data centers are contrasted between nine months to 18 months for gigawatt liquid-cooled facilities and significantly faster deployment for SambaNova equipment, which supports 10 kilowatt per rack configurations suitable for existing downtown spaces in cities like Paris or Manhattan.
  • Ultra-low latency requirements for agentic AI are predicted to demand response times under 0.1 seconds per agent to keep total interaction delays within 1–2 seconds, driving hardware deployment to large metropolitan areas and edge locations.
  • The market is expected to bifurcate into large gigawatt data centers costing $50–100 billion and a wave of mid-sized, distributed, modular deployments, such as shipping container clusters powered by solar or standalone connections.
  • Technology adoption will prioritize premium inference defined as running largest models at original precision without quantization, with model sizes projected to grow from current 1–2 trillion parameter open-source models toward 10 trillion parameters.
  • Infrastructure consolidation is predicted to result in only two to four chip types globally, with service providers differentiating through blended technologies, ultra-low latency for specific agents, and secure data-private inferencing.
  • Business strategies are expected to shift from cost saving to revenue generation via custom-trained models, with enterprise differentiation relying on specific data and IP rather than commodity models within an approximate two-year timeframe.
  • Global sovereignty trends are forecast to increase, with countries like Japan and Korea investing in homegrown models and enterprises repatriating infrastructure to on-prem solutions for data privacy and security.
  • The company intends to focus on shipping racks and software while partnering with firms like Cambium and Vista Equity for NeoCloud services, aiming to route traffic to SambaNova racks for lower cost and higher performance.
  • Competitive risks include the potential for companies failing to embrace scaling technology to lose market ground, while dominant players are expected to be those that successfully utilize AI for differentiation and global service scaling amidst worsening energy and cost constraints.