newsfilter.io
Fireside Chat, Interview

High Performance Compute: Building the Cloud for AI Teams | Erik Bernhardsson | RAISE Summit 2026

  • Modal projects its GPU infrastructure will reach 50,000 units by the end of the year, with a surge in new GPUs and CPUs expected over the next three months, while the broader GPU supply shortage is anticipated to remain tight for another couple of years before eventually catching up with demand.
  • The Sandbox product is predicted to have generated a significant revenue explosion last summer after a two-year period of low usage, while audio models like speech-to-speech are currently underperforming but expected to achieve breakthroughs and become "huge" within one to two years.
  • Eric plans to launch a dedicated CI/CD solution and an AutoEndpoints LM inference product that will dynamically improve via speculator models trained on traffic data, using agents in the loop to optimize hyperparameters and address bottlenecks in testing and code verification.
  • Modal intends to overcome GPU capacity constraints by sourcing independently across 20 different clouds and potentially constructing data centers, aiming to provide customers with access to 1,000 GPUs within seconds or minutes while managing routing, load balancing, and containerization.
  • The company plans to build a custom file system, container runtime, and scheduler to replace insufficient tools like Kubernetes and Docker, ensuring a fast, bi-directional streaming experience via a dedicated inference CDN, with 1,000 engineers identified as a persistent structural constraint similar to AWS.
  • Market trends indicate that companies will shift from frontier models like Anthropic to open-source models such as Qwen, Kimi, and GLM5 upon reaching commercial maturity for cost control, with DeepSeek having fundamentally altered the landscape for inference and Chinese models reaching parity with frontier counterparts.
  • Eric forecasts that open-source models utilizing proprietary data or reinforcement learning will eventually outperform frontier models and advises against building specifically for Europe due to regulatory headwinds, tax systems, and a perceived lack of "megalomaniac vision" compared to US founders.
  • Modal plans to open a new office in London soon while maintaining that the US will remain ahead on various fronts, with European growth hindered by boring regulatory issues rather than talent availability.
  • Customers are expected to be relieved of GPU supply concerns as Modal manages infrastructure complexities, including capacity and sourcing, allowing them to focus on fine-tuning models and adopting complex techniques like reinforcement learning once they reach a certain stage of commercial maturity.
  • Product prioritization will rely on a combination of personal experience building AI applications, customer feedback, and speculation regarding future industry trends, as Nvidia's dominance faces increasing competition and hyperscalers remain unable to secure sufficient capacity for Modal's needs.