newsfilter.io
Interview, Fireside Chat

Nvidia CTO Michael Kagan: Scaling Beyond Moore's Law to Million-GPU Clusters

  • Computing performance requirements are projected to grow 10x to 16x annually to support AI workloads where model sizes double every three months, with new products introduced yearly to deliver an order of magnitude higher performance per unit.
  • Data centers are expected to scale from 100,000 GPUs to potential clusters of one million, driven by the fusion of von Neumann machines and accelerated computing, with future facilities reaching gigawatt-scale power levels and discussions underway for 10-gigawatt sites.
  • The market is anticipated to expand exponentially as AI adoption shifts human behavior to enable 100x more work, with inference workloads becoming increasingly compute and memory-intensive, potentially exceeding training demand due to recursive generation and reasoning capabilities.
  • Physical and environmental constraints are expected to limit growth, specifically regarding the time required for concrete stabilization, heat management necessitating a shift to liquid cooling, and physics limits that end the era of density increases via Moore's Law.
  • System scaling will rely on networking capabilities to connect vast numbers of chips into a single fabric, addressing challenges such as speed-of-light latency variance across distant data centers and the need for narrow latency distribution to prevent jitter.
  • Future architectures will utilize Bluefield DPUs to isolate infrastructure computing on separate platforms for security, while specific GPU SKUs will be optimized for distinct pre-fill (compute-intensive) and decode (memory-intensive) tasks to allow flexible data center configurations.
  • The partnership between NVIDIA and Intel is expected to create a win-win market expansion, leveraging x86 architecture for general-purpose computing and accelerating it with specialized hardware to serve challenging markets.
  • Software will play a critical role in managing system-level challenges, including designing around hardware component failures, optimizing job placement across millions of components, and utilizing telemetry from devices like Spectrum X to manage network congestion without large buffers.
  • Historical and climate simulation tools like "Earth 2" are expected to enable experimental science capable of predicting global warming impacts 50 years into the future, with AI potentially discovering physical laws unknown to humans.
  • A new campus is planned for Israel to support regional growth, reflecting the successful cultural integration of the Mellanox acquisition which contributed to over 2x manpower growth in the region and significant market cap expansion.
  • Performance growth is projected to maintain a slope of 10x or several orders of magnitude per year, though the speaker notes the limits of this trajectory are difficult to define, with the industry shifting from "squeezing transistors" to "bringing together thousands of chips."
  • Customers are expected to optimize data centers by selecting specific SKUs for different workload phases, while the "single unit of computing" architecture will evolve to encompass millions of components connected by networking to maintain efficiency despite individual component failure rates.
  • The Ethernet business, alongside NVLink and InfiniBand, is expected to be the fastest-growing segment, providing the critical communication fabric necessary to scale beyond single-node capabilities and hide communication overhead behind computation.