Conference Presentation, Keynote
From Metal to Model: Why Operations Win the AI Infrastructure Race | Mirantis | RAISE Summit 2026
- Future IT work will integrate AI directly into workflows, with AI inference consuming two-thirds of compute power globally and rising to 75% by 2030.
- Infrastructure requirements are projected to increase rather than decrease due to efficiency gains, described as "Yevon's paradox," as industry clusters expand from 1,000 to between 8,000 and 25,000 GPUs.
- Failure rates in larger environments are expected to accelerate from one GPU failure every three hours to one every 30 minutes, with optics overheating potentially causing failures within 24 to 48 hours.
- Human intervention will remain necessary for approximately one out of every 133 failures (3 out of 400), while CPU resource needs will grow significantly to support intelligence workloads alongside GPUs.
- Hardware commoditization is expected to shift the durable advantage to platform, process, and full-stack ownership, as model architectures evolve faster than current infrastructure can accommodate.
- Customer choices regarding models and workloads are anticipated to move faster than provider capabilities, necessitating reserve capacity to maintain user experience despite new user influxes.
- Implementing specific processes and procedures is expected to increase uptime by nearly 40% or more, with the "Cordon AI" product positioned as the future standard for cloud infrastructure.
- A strategic plan involves building an open, standard-driven platform with open APIs, contrasting with "black box" systems where deep insights are absent and failure is deemed inevitable.
- Financial growth for future cloud providers is expected to depend on full-stack ownership to manage the complexities of evolving model architectures and inference approaches.