newsfilter.io
Fireside Chat, Conference Presentation, Panel

Operating AI Infrastructure at Multi-Model, Multi-Cloud Scale | RAISE Summit 2026

  • Inference spending is projected to expand at a rate comparable to broader AI growth, driven by customer demand for metrics like tokens, requests per second, and traffic patterns rather than hardware specifications.
  • Business operations may require adaptation to accommodate the industry's shift toward inference, with the company planning to deepen investments in model performance as more models are deployed.
  • To secure global compute capacity, the organization aims to aggregate demand from numerous AI-native companies while maintaining a cloud infrastructure spanning 18 to 20 providers across over 90 regions.
  • Technical capabilities will focus on scaling models across multiple regions to satisfy data residency, performance, and latency constraints, alongside optimizing the cost versus performance equation within the token factory.
  • Hardware execution will leverage internal expertise to plan for specific production debugging needs down to the physical PCB or cooler level, allowing for flexible design depth.
  • Future infrastructure needs regarding the quantity of data centers are uncertain, with estimates ranging from thousands to hundreds depending on market distribution between providers.
  • The market is expected to remain sufficiently large for all participants given the rapid industry growth, fostering an environment where competitors may also collaborate.
  • Open source models are anticipated to persist despite geopolitical tensions, with a gap between open source and frontier models expected to close through improvements in token efficiency and specialization.
  • Cost advantages in non-frontier models, such as lower cost per token, may be offset by reduced task performance and efficiency unless models undergo specific specialization to improve token usage.
  • Specialization efforts are predicted to enhance token efficiency, thereby reducing costs and enabling faster response times for novel user experiences.