Fireside Chat, Interview
Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance
Market & Architectural Realities
- Current GPU inference utilization is estimated at only 5–7%, meaning 93–95% of hardware capacity is wasted.
- The fundamental GPU architecture, relying on off-chip memory, is identified as a critical bottleneck for inference efficiency.
- Cerebrus predicts a reduction in dependence on Transformer architectures within three to five years due to inherent quadratic attention limitations.
- AI is shifting from a "novelty" phase (late 2024) to a "utility" phase, integrating into daily workflows for non-tech demographics.
- Inference market volume is expected to grow over 100x in the next five years driven by increased user bases, frequency of use, and compute-per-instance requirements.
Cerebrus Hardware & Wafer-Scale Innovation
- Cerebrus utilizes wafer-scale computing to replace slow HBM with massive on-chip SRAM, enabling high-capacity storage of 70B+ parameter models without the data movement latency of traditional chips.
- The architecture uses hundreds of thousands of identical silicon tiles with redundancy, allowing defective tiles to be disabled and bypassed rather than discarding the entire chip.
- This yield-management technique, applied for the first time in 70 years of chip history, allows Cerebrus to produce whole-function wafers despite natural silicon flaws.
- Inference speed is prioritized over training speed, as milliseconds of delay in interactive mode destroy user attention and business viability.
- Power efficiency is achieved by keeping data movement within the silicon domain, reducing the high power consumption associated with off-chip IOs.
Industry Trends & Strategic Outlook
- Synthetic data is projected to comprise nearly 100% of training data within five years, used specifically to simulate rare, high-value scenarios like pilot emergencies or complex medical procedures.
- Algorithmic efficiency is expected to improve significantly as models move from "all-to-all" connections to sparse, state-based, or mixture-of-experts (MoE) architectures.
- Cerebrus remains the sole provider of wafer-scale AI hardware, creating a unique market position compared to NVIDIA, TPU, or Cerebras.
- CUDA lock-in is dismissed as non-existent for inference workloads; PyTorch compilation allows users to switch hardware vendors with minimal friction.
- NVIDIA is predicted to retain 50–60% market share in five years, but their dominance will face pressure from customer delays and architectural limitations.
- Enterprise value will likely favor chip providers over model providers long-term, as hardware moats are harder to replicate than software models.
Operations, Geopolitics & Ethics
- G42 represents 87% of Cerebrus's revenue, serving as a strategic learning platform to scale manufacturing and supply chain readiness for future hyperscaler partnerships.
- Cerebrus has voluntarily refused to sell to China due to concerns regarding the potential military application of the technology and use in human rights violations.
- Export controls on hardware are deemed more enforceable than software controls due to the physical weight, tracking requirements, and lower risk of diffusion.
- The company maintains cash flow positivity and positive gross margins, distinguishing it from the hemorrhaging cash models of many software-only AI competitors.
- The U.S. faces infrastructure bottlenecks where power exists (e.g., Niagara) but lacks fiber connectivity and regulatory approval for data center siting in populated regions.
- China's AI capabilities are not underestimated, citing their exceptional infrastructure investment, state-backed industrial policy, and engineering talent generation rates.
Leadership & Future Vision
- CEO Andrew Kaspar admits to being wrong on the initial rejection of water cooling, a decision he now acknowledges as critical as it is now standard in industry.
- Kaspar advocates for serial entrepreneurship in hardware, arguing that experience in supply chain management and large-scale engineering is essential, unlike in social media startups.
- Cerebrus aims to solve two major societal problems (e.g., finding therapeutic cures) and power a suite of undiscovered applications within 3–5 years.
- AI penetration is predicted to reach cell phone levels within two years, with 90% of code potentially written by machines within one year.
- The "modest" business focus in Arab states (UAE, Qatar, KSA) is seen as a driver for regional stability and closer integration with Western economies.