Interview, Podcast
Chasing Silicon: The Race for GPUs
- AI hardware demand is currently 10 times higher than supply and expected to persist due to exponential growth, with supply unable to react quickly as new semiconductor fabrication plants require a couple of years and investments ranging from a couple of billion to 10 billion to build.
- Companies will likely need to pre-reserve capacity, often securing exclusive access for two years, or seek investment deals with cloud providers, while founders may shop around for specialized clouds rather than large ones.
- Renting capacity via cloud or SaaS is considered the preferred strategy for most founders unless they have specialized needs or geopolitical concerns; owning infrastructure is deemed viable only at a scale of approximately $100 million annually, whereas spending $10 million annually is considered too low for critical independent operations.
- Industry trends point toward using slightly smaller models trained more efficiently to match larger model performance based on Chinchilla scaling laws, with a future expectation of training on public data first followed by fine-tuning on private data, while large open-source LLMs comparable to GPT-3 (175 billion parameters) are not yet available and the gap is expected to persist.
- As devices become faster and models more optimized, basic large language and image generation models are expected to integrate into operating systems, creating a bifurcation where simple tasks run locally while complex tasks remain in the cloud.
- A "Cambrian explosion" of creativity is forming within a new ecosystem where software becomes increasingly important relative to hardware, presenting opportunities to build new application stacks as data generation continues to grow.
- Future discussions are anticipated to cover how compute costs will change over time and the sustainability of current AI spending levels.