Fireside Chat, Conference Presentation, Panel
Operating AI Infrastructure at Multi-Model, Multi-Cloud Scale | RAISE Summit 2026
- Inference spending is projected to expand at a rate comparable to broader AI growth, driven by customer demand for metrics like tokens, requests per second, and traffic patterns rather than hardware specifications.
- Business operations may require adaptation to accommodate the industry's shift toward inference, with the company planning to deepen investments in model performance as more models are deployed.
- To secure global compute capacity, the organization aims to aggregate demand from numerous AI-native companies while maintaining a cloud infrastructure spanning 18 to 20 providers across over 90 regions.
- Technical capabilities will focus on scaling models across multiple regions to satisfy data residency, performance, and latency constraints, alongside optimizing the cost versus performance equation within the token factory.
- Hardware execution will leverage internal expertise to plan for specific production debugging needs down to the physical PCB or cooler level, allowing for flexible design depth.
- Future infrastructure needs regarding the quantity of data centers are uncertain, with estimates ranging from thousands to hundreds depending on market distribution between providers.
- The market is expected to remain sufficiently large for all participants given the rapid industry growth, fostering an environment where competitors may also collaborate.
- Open source models are anticipated to persist despite geopolitical tensions, with a gap between open source and frontier models expected to close through improvements in token efficiency and specialization.
- Cost advantages in non-frontier models, such as lower cost per token, may be offset by reduced task performance and efficiency unless models undergo specific specialization to improve token usage.
- Specialization efforts are predicted to enhance token efficiency, thereby reducing costs and enabling faster response times for novel user experiences.