Interview, Other
The True Cost of Compute
- Training costs for large language models, currently estimated at tens of millions of dollars, may stabilize or slightly decline as the industry shifts from compute constraints to data scarcity, though absolute spending is expected to rise for the foreseeable future as the sector remains in its infancy.
- While the proportion of total expenses allocated to compute resources will likely decrease over time as companies prioritize product offerings and headcount grows, compute demand is not expected to subside soon, with access to resources serving as a critical success factor for AI firms.
- Current funding dynamics show that many companies allocate over 80% of raised capital to compute, where capital barriers are viewed as manageable speed bumps rather than insurmountable moats, fostering expectations for increased future innovation.
- Data availability limits rapid scaling, as sufficient knowledge for hundred-fold increases in model size or training data volume has not yet been produced, with specific training datasets ranging from hundreds of gigabytes to 100 billion characters for models like ChatGPT and 2 trillion tokens (approximately 8 trillion characters) for Lama 2.
- Hardware utilization rates can be improved to the tens of percent range through optimization, contrasting with naively implemented systems that often remain below 10%.
- Renting an A100 chip typically costs between one and four dollars, but securing this capacity often requires a two-year reservation despite actual usage periods of only two months.
- Inference expenses for large language models are estimated at a fraction of a cent (roughly a tenth to hundreds of a cent), and image models may run on consumer-grade hardware like MacBook graphics cards or local consumer GPUs to reduce costs.
- Training efficiency relies heavily on high-performance chips, as the overhead of distributing data across less performant processors likely negates any cost savings from using cheaper hardware.