Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Infrastructure, Capstone Case
Intel's market position is currently turning due to two primary tailwinds: global supply constraints favoring companies with domestic manufacturing capabilities, and a resurgence in CPU utility driven by the rise of AI agents.
Intel remains the only leading-edge American company with significant in-house manufacturing capacity, a factor that provides a distinct competitive advantage in the current hardware-constrained environment.
Kati Katti joined OpenAI in November after serving as Intel's CTO and head of AI, a transition he notes coincided with the market finally recognizing Intel's AI transformation.
OpenAI is targeting 30 gigawatts of total compute capacity by the end of the decade, a figure that represents a massive scaling of its current infrastructure.
Over the last three years, OpenAI has tripled its compute capacity annually, a trend that is hyper-correlated with its revenue growth.
Revenue is identified as a lagging indicator for frontier labs, primarily driven by the volume of available compute, its utilization rate, the number of users, and total token consumption.
As of late 2024, the model Codex has seen double-digit growth in just two weeks following the release of version 5.5, with usage expanding beyond coding to general-purpose knowledge work.
OpenAI's mission prioritizes maximizing compute availability for unconstrained research exploration over short-term revenue maximization.
The compute workload split is shifting decisively toward inference, with projections indicating that inference will comprise over 80% of total compute usage in the future.
Scaling laws have evolved to apply across the entire AI lifecycle, including pre-training, post-training (RL), and synthetic data generation, which are primarily inference-heavy workloads.
OpenAI's strategic goal regarding token economics is threefold: reduce the cost per token via hardware/software improvements, increase token intelligence, and reduce the total number of tokens required per task.
Securing compute at the gigawatt scale involves sourcing the entire supply chain, including chips, memory, networking, power generation, cooling, land, and data center construction.
The primary operational bottleneck for Katti's team is not signing contracts but orchestrating the delivery and integration of millions of chips to ensure they function reliably at scale.
A gigawatt of compute requires approximately 500,000 GPUs, demanding precise engineering to manage power distribution, cooling, and system stability against fluctuations.
Deploying gigawatt-scale data centers creates significant grid stability risks, potentially causing state-wide blackouts if synchronized training jobs create massive, rapid energy demand spikes.
Infrastructure decisions are being made to decouple AI power needs from the standard grid, utilizing natural gas and increasingly nuclear energy to ensure reliability.
The long-term vision includes a future where every human possesses a personal GPU (1–2 kW), requiring roughly 7 terawatts of global compute capacity, far exceeding current projections.
Hyperscalers collectively are planning to build approximately 100 gigawatts of compute, which would consume a double-digit percentage of total US energy capacity.
OpenAI's "compute advantage" manifests as the ability to serve models like o1 (5.5) without token limits or artificial caps, allowing for broader and more intensive user interaction.
Modern agentic workloads have evolved from simple one-shot inference to complex, multi-step processes that close the loop by trying actions, iterating, and refining outputs.
The compute graph for agents has shifted from a simple node structure to a directed acyclic graph involving multiple inference calls, tool usage, database queries, and virtual machine spinning.
Current user experience bottlenecks are AI latency and execution time; the goal is to make AI so fast that the human decision-maker becomes the primary bottleneck.
Delivering a seamless agentic experience requires heterogeneous compute infrastructure, as pure GPU setups cannot economically or efficiently handle the diverse needs of reasoning, tool use, and long-context retention.
Cerebras chips are being utilized to accelerate token generation, which forces OpenAI to optimize the rest of the software stack to eliminate latency from other layers.
The first token generation latency for models like Codex currently ranges from 400 to 500 milliseconds, primarily driven by the "pre-fill" phase where the model processes hundreds of megabytes of context (e.g., 400k tokens).
Future infrastructure may move toward specialized accelerators for long context (e.g., holding entire codebases) and fast inference, rather than relying on a single chip type.
The AI supply chain is expected to diversify beyond NVIDIA, driven by TSMC's strategy to allocate wafers across multiple customers to de-risk its business.
Concentrated, large-scale compute clusters remain economically superior to distributed edge clusters due to labor scarcity, the high cost of small-scale buildouts, and the current inability of models to be sufficiently distilled for low-latency edge execution.
Katti identifies the most misunderstood aspect of AI infrastructure as the current reliance on simplistic system designs; he predicts a future where AI models recursively design the very chips and software architectures needed to run them.
The industry is moving away from a three-year chip design cycle to a faster, AI-driven iteration loop to keep pace with model advancements.
Historically, profitability in tech revolutions shifts from infrastructure to platforms and apps; Katti predicts the AI profit wave will follow a similar trajectory, eventually moving up the stack.
Katti is "long" on the lowest layers of the AI stack, specifically foundational hardware infrastructure like transformers, batteries, cooling, and component manufacturing, citing these as underserved and high-barrier-to-entry.
He is "short" on model wrappers and traditional app-based interfaces, arguing that the AI ecosystem will evolve toward outcome-based interaction where users define results rather than selecting specific applications.
Katti identifies the single largest structural issue in current infrastructure as the lack of sufficient fab capacity for logic and memory, creating a critical choke point across TSMC, Samsung, Intel, and ASML.
In the short to medium term, software innovation is expected to focus on orchestration, harnessing, and token efficiency.
In the medium to long term, the most significant software breakthroughs are expected to come from new memory architectures surrounding compute units, as the transformer architecture itself is unlikely to change fundamentally soon.
OpenAI maintains that open-weight models will remain six months behind frontier models, justifying continued massive investment in proprietary, large-scale intelligence.
For the StarGate project, OpenAI prioritizes "time to compute" over raw capacity, favoring the construction of large, concentrated gigawatt sites over fragmented smaller deployments to ensure rapid operational readiness.