Interview, Fireside Chat
Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China
NVIDIA-Intel Strategic Partnership and Market Dynamics
- NVIDIA announced a $5 billion strategic investment in Intel to jointly develop custom data centers and PC products.
- The market reacted immediately, with NVIDIA stock rising approximately 30% following the announcement.
- The deal involves Intel creating a chiplet to be packaged alongside NVIDIA GPU chiplets for PC integration.
- Analysts characterize the partnership as a "Warren Buffett effect," signaling high confidence in the semiconductor giant's future.
- The collaboration reverses a historical dynamic where Intel previously sued NVIDIA for anti-competitive behavior regarding chipset graphics integration.
- Intel is now "crawling" to NVIDIA, acknowledging the shift in power from x86 CPU dominance to GPU-centric computing.
- A fully integrated x86 laptop with NVIDIA graphics is described as potentially the best market product, surpassing ARM-based alternatives for specific use cases.
- Competitors face significant headwinds from this alliance:
- AMD is described as having received "worst possible news," compounding existing struggles with software stack traction against NVIDIA.
- ARM's value proposition is threatened as NVIDIA gains direct access to Intel's x86 architecture and legacy technologies.
- Capital injection details regarding Intel's survival strategy:
- NVIDIA's $5 billion and SoftBank's $2 billion investments are viewed as "small" relative to Intel's estimated $50 billion capital requirement.
- A $10 billion government commitment was also noted, though total capital needs may still require future public market dilution.
China's Semiconductor Self-Sufficiency and Huawei
- Huawei unveiled an AI roadmap in 2025, aiming to replace NVIDIA chips domestically following US export bans.
- Huawei is developing custom memory solutions, specifically splitting inference hardware into chips optimized for "pre-fill" and "decode" workloads.
- This mirrors NVIDIA's recent hardware segmentation but highlights China's struggle with High Bandwidth Memory (HBM) manufacturing capacity.
- Historical context on Huawei's resilience:
- By late 2020, Huawei had shipped seven-nanometer AI chips and briefly surpassed Apple as TSMC's largest customer before US bans took effect.
- In 2024, Huawei was found to have smuggled approximately 2.9 million chips through shell companies, valued at roughly $500 million.
- Supply chain constraints for China remain critical:
- While logic chip manufacturing (7nm and potentially 5nm) is ramping, HBM production is a major bottleneck requiring specialized etch equipment imports.
- China's import data shows a surge in etch equipment acquisitions, a necessary step for stacking HBM wafers, though yield rates lag behind Western leaders.
- Domestic demand for high-performance chips is currently outpacing production, creating a transition gap where companies like ByteDance still prefer NVIDIA hardware.
- Strategic implications for US policy:
- China's ban on NVIDIA may be a negotiation tactic to force the US to relax export controls, leveraging the narrative of domestic readiness.
- Smuggling of NVIDIA chips via third countries continues, suggesting a continued reliance on foreign hardware despite official mandates.
- The "Galapagos" theory is discussed: isolating China's tech ecosystem could force them to develop unique, non-global technologies, potentially allowing them to dominate non-US markets (Middle East, Southeast Asia) with local alternatives.
NVIDIA Moat, Strategy, and Financial Outlook
- Jensen Huang's leadership relies on "gut instinct" rather than traditional spreadsheet forecasting, often ordering capacity months before receiving confirmed customer orders.
- This strategy has led to billions in inventory write-downs during past downturns (e.g., crypto mining) but allowed NVIDIA to dominate when demand surged.
- NVIDIA consistently ships "A0" (first revision) silicon without major stepping errors, a rarity in the semiconductor industry that accelerates time-to-market compared to AMD or Intel.
- Bull case for NVIDIA's future growth:
- Hyperscaler CapEx consensus is $360 billion for next year, but independent estimates suggest $450–$500 billion.
- The long-term vision includes AI infrastructure spanning billions of nodes, potentially capturing a $2 trillion market value.
- AI "takeoff" scenarios where AI builds more AI could exponentially increase compute demand beyond current linear projections.
- Challenges regarding NVIDIA's massive cash reserves:
- The company generates hundreds of billions in free cash flow annually, yet faces regulatory hurdles preventing large-scale acquisitions (e.g., ARM).
- Jensen Huang is unlikely to directly compete in cloud infrastructure or buy entire startups, fearing customer alienation.
- Potential capital deployment areas include investing in data center power generation, energy infrastructure, and providing backstop financing for niche cloud providers (e.g., CoreWeave).
- Risk management strategy involves "betting the farm" on supply chain expansion before revenue is realized, a high-risk, high-reward model that has historically paid off.
Hyperscaler Performance and Cloud Infrastructure
- Amazon Web Services (AWS) is projected to see revenue re-acceleration to over 20% year-over-year following a trough in early 2024.
- AWS possesses the largest inventory of spare data center capacity globally, which will be repurposed for AI workloads using Tranium and NVIDIA GPUs.
- Despite structural inefficiencies in cooling and networking (Elastic Fabric), the sheer volume of physical capacity allows AWS to compete on scale.
- Oracle is identified as a major winner in the AI compute market:
- Oracle secured a multi-year, $300 billion deal with OpenAI, leveraging its balance sheet and flexibility to host non-NVIDIA networking (Arista, Broadcom) alongside NVIDIA hardware.
- Independent analysis tracked Oracle's data center signings and power procurement to predict revenue growth, validating the deal's scale.
- Microsoft is shifting from an exclusive compute provider to a "write-a-first-refusal" model, opening the door for Oracle to capture OpenAI's massive workload.
- Infrastructure evolution includes:
- Shift toward gigawatt-scale data centers (e.g., XAI's Colossus) as 100kW clusters are now considered standard.
- Rapid construction timelines (6 months) achieved by leveraging cross-border power solutions (e.g., Elon Musk moving power plants across state lines).
Hardware Trends: Blackwell, Inference Optimization, and Market Cycles
- The GPU market has transitioned from a pure capacity crunch to a nuanced environment where deployment complexity drives demand.
- NVIDIA's Blackwell (GB200) faces deployment delays due to reliability challenges with its 72-GPU NVL 72 form factor.
- A single GPU failure in a 72-GPU rack can take the entire node offline, necessitating complex workload splitting between high-priority and low-priority tasks.
- Total Cost of Ownership (TCO) for GB200 is estimated at 1.6x the H100, but performance gains can range from 2x to 6x depending on the workload.
- AI chip architecture is bifurcating to optimize inference:
- Hardware is splitting into dedicated "pre-fill" chips (optimized for FLOPs on long context) and "decode" chips (optimized for memory bandwidth on token generation).
- This disaggregation allows providers to auto-scale resources based on user traffic patterns (short input/long output vs. long input/short output).
- Startups like Cerebras and new NVIDIA products (CPX, DPX) are targeting these specific workloads to improve efficiency.
- Current market status:
- The "buying GPUs like cocaine" analogy describes the current informal, high-friction market where capacity is secured via direct communication rather than formal RFPs.
- Prices for Hopper (H100) GPUs have bottomed and are beginning to rise as supply tightens and Blackwell deployment ramps.
- The industry is in a "growing pain" phase where the learning curve for Blackwell infrastructure is slowing the effective supply of high-performance compute.