Interview, Fireside Chat
Steeve Morin: Why Google Will Win the AI Arms Race & OpenAI Will Not | E1262
- Compute demand is projected to shift such that in five years, 95% of resources will support inference while only 5% remains for training, driven by agents and reasoning tasks requiring low-latency single-stream performance rather than aggregate throughput.
- The market will transition from treating models as single abstractions to constellations of backends where APIs serve as the primary interface, enabling switching between hardware providers to achieve up to an order of magnitude efficiency gain.
- Specific hardware migration strategies, such as moving from NVIDIA to AMD for a 70B model, are expected to yield four times better spending efficiency, though reliance on HBM and NVIDIA's supply chain dominance may create competitive barriers.
- Current GPU financial models, including six-to-seven-year amortization plans, are viewed as unsustainable and potentially part of a bubble that will burst due to price increases outpacing inference performance gains, leading to potential distressed sales at 30 cents on the dollar.
- An oversupply of purchased compute is anticipated within the current year as hyperscalers face overbuying risks, potentially resulting in massive overhangs where chips become outdated before deployment.
- Industry incentives will shift toward "less is better" inference paradigms prioritizing reliability and avoiding overnight maintenance, with auto-scaling predicted to offer 5x to 10x savings over always-on provisioning.
- Competition will intensify across infrastructure layers, favoring a mix of GPU, TPU, and dedicated ASIC options from providers like Google and Amazon, which may force margins for cloud providers and chip sellers to decrease.
- New competitors including Cerebras, Etched, Vsor, Rain, and Fractile are expected to challenge NVIDIA by offering competitive pricing, "compute in memory" architectures, and lower unit costs for high-performance inference.
- Talent and energy availability are identified as the primary limiting factors for AI scaling, outweighing capital expenditures such as data center projects like Stargate.
- Structural changes include a move away from continuous fine-tuning toward smaller, efficient models and runtime specialization, alongside the potential obsolescence of current transformer architectures by non-transformer models like JEPA.
- Synthetic data usage will standardize in verticals like code generation where execution validates output, though general injection may cause model deterioration, while export restrictions in regions like China continue to drive constraint-based architectural innovation.
- AI startups may struggle if relying on reselling compute for margins, as 98% of spend could go toward upstream hardware margins, while NVIDIA faces potential demand downslopes due to Blackwell chip issues and order cancellations.
- Future compute architectures will likely integrate more SRAM directly onto chips, and market adoption will increasingly favor any provider avoiding the 90% margins associated with NVIDIA chips despite potential software maturity gaps.