Interview, Fireside Chat
Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis
- Semi-Analysis employs 90 staff with projected growth to support scaling, aiming for revenue at or near $100 million and potentially launching a venture fund driven by ecosystem demand.
- Inference is predicted to surpass oil as the largest global market, accounting for multiple percentage points of global GDP, while inference model costs for equivalent quality are forecast to drop 60x annually.
- The InferenceX project intends to utilize over $50 million in donated hardware immediately, reaching over $100 million with TPUs, and will run daily benchmarks on leading models from Moonshot, Alibaba, Chinese labs, and US open-source groups.
- AI infrastructure is expected to bifurcate into optimized batch or instant response systems, diverging from a "one size fits all" model as future architectures require specific hardware co-optimization.
- By 2030, OpenAI and Anthropic combined are projected to consume over 100 gigawatts, with global AI compute consumption reaching terawatt levels by 2040.
- Compute capacity located in space is expected to exceed 50% of incremental additions after 2030, though sub-1% of total capacity will reside in space by 2030.
- Efficiency metrics include intelligence per watt improving at a rate of approximately 40x annually and memory capacity/bandwidth accelerating through direct-on-chip stacking.
- Chip power density is forecast to exceed one watt per millimeter squared, enabling 4000-watt chips and reduced silicon footprints despite thermal challenges, with US capacity to convert millions of diesel engines to generators for data center power.
- Production targets for 2025 include over 10 million Google TPUs and tens of millions of NVIDIA GPUs, with revenue projections exceeding $100 billion annually for TPUs and $500 billion for GPUs.
- Major labs and hyperscalers are expected to pursue private ASIC programs, deploying billions to hundreds of billions in custom silicon annually, while the AI total addressable market expands faster than compute capacity can be built.
- Anthropic is expected to achieve net income profitability by Q3, driven by gross margins on OpenAI's Opus token exceeding 80%, while Google and Meta capital spending continues to rise rapidly.
- Data center power rental rates are forecast to remain between $120 and $200 per kilowatt-month, with specialized operators charging premiums for facilities operating at 1.5x nameplate capacity.
- NVIDIA is expected to fund "neo-labs" and "neo-clouds" to maintain a multipolar ecosystem, while co-packaged optics are anticipated to become the industry standard between 2027 and 2030.
- The "CUDA moat" is predicted to erode as models gain the capability to write their own optimized kernels, though DeepMind and certain Google projects may continue relying on NVIDIA GPUs for specific science applications.