Interview, Fireside Chat, Conference Presentation
Poolside AI, IREN, Forbes: From 10x to 100x Building AI Systems for Real World Scale
Key Participants and Strategic Focus
- Dennis Skriniop serves as CTO of Iron, a data center provider specializing in sustainable, large-scale infrastructure for next-generation AI workloads.
- Isa Kent is Co-founder and CTO of Poolside, a foundation model company building software development capabilities oriented toward human-level intelligence.
- Poolside has raised $500 million to pursue AGI, using software development as the primary proxy for measuring intelligence.
Scaling Foundations and Operational Complexity
- Compute Requirements: The trajectory toward human-level intelligence requires an order of magnitude increase beyond current clusters of ~10,000 H200 GPUs.
- Scaling Axes: Model intelligence growth is driven by three variables: model size, fossil fuel data (web), and renewable energy data (reinforcement learning).
- Experimental Velocity: Poolside's 40-person applied research group executed ~1,800 experiments in June alone, necessitating a "model factory" approach rather than artisanal engineering.
- Engineering Ratio: Approximately 80% of foundation model building is attributed to hardcore distributed systems engineering and scaling automation rather than pure algorithmic discovery.
- Infrastructure Specialization: Clusters exceeding 1,000 GPUs and 10,000 CPUs require specialized roles including liquid-to-chip engineers, mechanical engineers, HPC DevOps, and InfiniBand experts.
- Failure Management: As scale increases, failure rates and infrastructure optimization costs rise exponentially, making time-to-resolution critical.
Technical Architecture and AGI Measurement
- Training vs. Inference: Model plasticity and generalization capabilities decrease the further into training a model is; thus, early pre-training is critical for unlocking complex reasoning.
- Three Optimization Pillars: Continuous improvement requires optimizing (1) compute efficiency (hardware to architecture), (2) data quality and efficiency, and (3) research breakthroughs via novel data.
- Evaluation Cadence: Every experiment triggers automatic evaluations 10 to 100 times during its lifecycle, moving beyond flawed individual benchmarks to spectrum analysis of reasoning and task capabilities.
- Compute Reality Check: While $500 million is a significant raise, it represents insufficient runway for training at scale over a year; the industry must prioritize scaling training compute over marketing metrics.
- Infrastructure Agnosticism: Poolside has adopted a strategy of disaggregating data storage from GPU clusters, streaming data from external zones (e.g., nearest wired-up AWS zones) to prevent I/O bottlenecks.
- Automation: Large training runs are fully automated for restarts and observability; however, experimentation remains complex, requiring dynamic resource allocation and resizing.
Infrastructure Planning and Risk Management
- Client Demand Profile: Clients typically request massive clusters with only 30 days of lead time, forcing providers to build for "optionality" and procure hardware 6–9 months in advance.
- Primary Failure Point: Scaling from 200 to 10,000 GPUs introduces a "different mentality" regarding cooling, hot-spot management, cabling, and East-West network integration.
- Storage Strategy: While some storage is co-located for cost-effectiveness, the most effective paradigm places heavy data ingestion and CPU-based tasks outside the GPU cluster to maximize training throughput.
- Network Criticality: At scale, the network and storage platform integration become the primary bottlenecks; everything must work in unison to handle massive data flow.
Energy, Sustainability, and Market Dynamics
- Energy Mix Reality: Dennis asserts that relying solely on renewables is unrealistic for 24/7 peak capacity data centers; a mix including nuclear and natural gas is currently necessary.
- Future Energy Outlook: The industry anticipates a 10–20 year horizon where battery storage and SMRs (Small Modular Reactors) become viable, though current lead times for these projects are significant.
- Inference Dominance: While training is CapEx-heavy, the primary future energy demand will come from inference, driving the need for flexible, always-on power infrastructure.
- Hyperscaler vs. Neo-Cloud: Over 200 Neo-Clouds have emerged to fill the gap left by hyperscalers, offering ground-up, AI-specific infrastructure that is more cost-competitive and agile.
- Diverse Compute Strategy: Isa predicts a future "compute mix" where clients utilize a combination of hyperscalers, Neo-Clouds, and owned infrastructure, similar to the future energy mix.
- Agility Advantage: Specialized providers can bypass the slow decision-making cycles of hyperscalers to deliver cost-effective services tailored specifically to AI cluster requirements.