newsfilter.io
Interview, Fireside Chat

Dylan Patel — The single biggest bottleneck to scaling AI compute

  • The Big Four forecasted $600 billion combined CapEx for the current year, with portions of this spending allocated to turbine deposits for 2028–2029 and data center construction for 2027, reflecting a long-term commitment to compute capacity coming online over subsequent years.
  • OpenAI and Anthropic have secured $110 billion and $30 billion in funding respectively, providing sufficient capital to cover compute spend for the year without counting on earned revenue, while Anthropic aims to expand from roughly 2 gigawatts to over 5 gigawatts by year-end to support $60 billion in revenue growth over the next 10 months.
  • OpenAI is dominating capacity allocation by signing long-term deals for H100s at prices as high as $2.40, crowding out suppliers and securing compute at significantly higher margins than the $1.40 Hopper margin, while Anthropic relies on indirect access via Bedrock, Vertex, or Foundry models.
  • The semiconductor supply chain is projected to be the primary bottleneck for AI scaling through 2030, specifically regarding EUV tool manufacturing, with ASML's capacity capped at approximately 100 tools by the end of the decade, which limits total AI chip deployment to roughly 200 gigawatts despite a potential market demand for 52 gigawatts annually.
  • Memory availability is a critical constraint as the sector shifts from consumer devices to AI, with HBM4 stacks transferring 2.5 terabytes per second compared to DDR5's 64–128 gigabytes; smartphone volumes are expected to halve to 500–600 million units to free up wafer capacity, causing DRAM prices to rise more than NAND prices.
  • Long-term contracts provide significant margin advantages as prices for locked-in compute (2–5 years ago) diverge from current market rates, with the majority of the market currently locked in long-term agreements, while the depreciation cycle for GPUs could extend beyond five years if AI adoption remains high.
  • Infrastructure investment outcomes depend on AI timelines, with fast timelines favoring the US through high returns on capital and infrastructure scale, whereas slow timelines could allow China to catch up or surpass the US if it achieves fully indigenized EUV supply chains by 2030, despite current production limitations.
  • Power is not the primary constraint for the next decade, as it can be scaled through behind-the-meter turbines and grid unlocking, whereas labor constraints are being mitigated through modularization and factory-integrated blocks that reduce on-site personnel needs.
  • Geopolitical dynamics suggest that China's ability to manufacture at high volume using EUV tools by 2030 remains uncertain due to "production hell," with current capabilities estimated at 100 DUV tools annually compared to ASML's hundreds, while US leaders are aggressively securing 3-nanometer supply and energy assets.
  • Future chip performance improvements, such as the 20x inference advantage of Blackwell over Hopper, will be limited by networking speeds, memory bandwidth, and cooling capabilities within CoWoS packaging, while space data centers remain infeasible for the current decade due to deployment delays and cooling inefficiencies compared to Earth-based facilities.
  • Apple's share of TSMC's revenue is expected to decline as AI chip volumes (N2, A16) become dominant, potentially forcing Apple to pre-book capacity and prepay for CapEx two years out, while TSMC prioritizes allocation for AI chips over CPU business due to higher margins and stability.
  • Market participants generally believe current infrastructure spending numbers are accurate or underestimated, driving hedge funds to trade on expectations of a memory crunch, with model vendors expected to see margin increases due to severe capacity constraints and the necessity to destroy demand.
  • Architectural shifts such as NVIDIA's rack-scale all-to-all topology and Amazon's dragonfly approach are emerging to address scaling limitations, whereas Google's TPU fleet allows for larger production models and faster RL feedback loops, though parameter scaling remains slow due to memory capacity requirements and the 5x rollout multiplier for larger models.