newsfilter.io
Conference Presentation, Panel

AI Has a Memory Problem | Gelsinger, SK hynix & More | RAISE Summit 2026

  • The Core "Memory Wall" Problem: AI development is currently bottlenecked by a trio of conflicting demands: insatiable demand for token economics, increasing requirements for larger context windows and model sizes, and the need for drastically lower first-token latency.
  • Current Industry Status: The semiconductor memory sector has shifted from a historically cyclical industry (one good year out of four) to a sustained boom driven by AI, with memory companies reporting margins as high as 90%.
  • HBM Limitations: High Bandwidth Memory (HBM) is characterized as a "lousy" or "lousey" solution that is bitwise, power, and bandwidth inefficient; specifically, for every bit created for HBM, four bits of standard memory capacity are effectively sacrificed.
  • HBM Structural Flaws:
    • Stacked DRAM creates a thermal "sandwich" issue, leading to overheating risks.
    • Bandwidth is limited by the point-to-point connection (shoreline) between memory and the GPU.
    • It remains the "best" available option only because it is the "tallest midget" in the current landscape.
  • Future Architectural Solutions:
    • Industry consensus points toward integrating memory directly onto or into compute layers (stacked memory) rather than horizontal placement.
    • Companies like Cerebras, D-Matrix, and Next Silicon are pioneering hybrid bonding and micro-bumping to bypass shoreline limitations.
    • These new architectures promise bandwidth improvements 10 to 50 times greater than current HBM solutions.
    • These breakthroughs are not expected to reach volume production until late in the decade.
  • Market Consolidation:
    • The number of memory players has consolidated from approximately 20 to three major entities: SK Hynix, Samsung, and Micron.
    • SK Hynix and Samsung collectively control roughly 70% of the HBM market, with Micron holding the remaining 30%.
    • Second-tier players (e.g., Nanya, PSMC, XMC) exist but face a significant gap in capability compared to the top three.
  • Supply Constraints and "HBM Squeeze":
    • HBM production is roughly four times less efficient than commodity DRAM regarding wafer output due to the need for larger die areas (3x size) to accommodate IO logic.
    • Vertical stacking of 12–16 layers compounds yield losses, further reducing throughput.
    • Manufacturing throughput is significantly lower; for example, if a waffle maker produces 12 units/hour normally, the HBM process reduces this to 4 units/hour.
  • Supply Chain Geopolitics:
    • Production is heavily concentrated in Asia: DRAM fabrication in South Korea, logic manufacturing in Taiwan (TSMC), and NAND in Japan.
    • A blockade of Taiwan for just three weeks could brown out the global industry, highlighting a lack of supply chain resilience.
    • New government initiatives (e.g., U.S. CHIPS Act, TerraFab) aim to build onshore capacity, though major factories take approximately four years to reach operation.
  • Economic Shift: From Commodity to Strategic Asset:
    • Memory is transitioning from a volatile commodity to a critical, non-discretionary infrastructure investment for AI.
    • Designs now involve deep co-optimization between logic and memory partners years in advance, making supplier switching difficult and reducing commoditization.
    • The industry is moving toward long-term, contract-based agreements (5+ years) rather than spot market purchasing.
  • Supply/Demand Dynamics:
    • The market is currently operating at a 26% deficit in HBM and a similar deficit in DRAM.
    • Apple's recent price increases across product lines signal expectations that this shortage will persist for the foreseeable future.
    • Years of "plenty" are estimated to be several years away; demand growth is exponential (driven by longer context windows and agent scaling), while supply growth is linear.
  • Investment Risks and Catalysts:
    • Bear Case: A "bullwhip effect" where manufacturers oversupply based on projections, or a Jevons paradox where improved model efficiency drastically reduces memory intensity per token before demand scales.
    • External Shocks: Potential disruptions from geopolitical unrest, power grid failures, or delays in data center construction could abruptly alter demand curves.
    • Price Elasticity: While AI servers are less price-sensitive, the 75% of the market comprising consumer electronics (PCs, phones) remains highly price-elastic, posing a risk for demand restriction if prices remain high.
  • Forward-Looking Statements:
    • Pat Gelsinger predicts that "it is never been a greater time to be a memory person" due to the fundamental shift in AI infrastructure needs.
    • HBM supply constraints will continue for several more years regardless of capital investment due to the complexity of designs.
    • The industry must fundamentally solve the memory architecture problem to realize the full benefits of AI, as computation alone is no longer the primary bottleneck.