Interview, Fireside Chat, Conference Presentation
Poolside AI, IREN, Forbes: From 10x to 100x Building AI Systems for Real World Scale
- Next-generation training clusters are expected to exceed the current scale of 10,000 H200s by more than an order of magnitude in the near future.
- Poolside plans to execute approximately 1,800 experiments monthly, a volume achievable only by engineering for scale rather than hiring.
- 80% of foundation model development is predicted to require advanced distributed systems engineering to manage failure rates and optimization challenges.
- Poolside has secured $500 million in funding, a figure anticipated to grow rapidly as demand for compute directly correlates with model quality.
- Human-level capabilities across most knowledge work performed on laptops are projected to emerge within two to three years.
- Inference demand is expected to reach immense levels, necessitating global infrastructure development where everyone builds out systems.
- Foundation model companies will likely adopt a "mix of compute" strategy sourcing from multiple providers as demand expands.
- Clients frequently request cluster availability within 30 days, compelling providers to maintain flexibility and build for optionality.
- Providers must initiate procurement and infrastructure orchestration with OEMs six to nine months in advance to meet urgent delivery windows.
- Renewable energy alone is currently deemed insufficient for 24-7 peak load operations, though a future mix including nuclear is anticipated within 10 to 20 years.
- Significant lead times for wind, solar, and Small Modular Reactors (SMRs) create current constraints that require resolution.
- Scaling from 200 to 10,000 GPUs will fundamentally break existing assumptions regarding staffing, cooling, and networking.
- Poolside intends to separate GPU training/inference from data streaming and CPU work to reduce materialization time from hours or days to minutes.
- Foundation model companies report achieving two to three times greater training and inference efficiency annually in a continuous optimization cycle.
- Model plasticity and generalization ability are expected to decrease as training progresses, requiring improvements in compute efficiency, data quality, and research breakthroughs.