newsfilter.io
Interview, Fireside Chat

Jonathan Ross, Founder & CEO @ Groq: NVIDIA vs Groq - The Future of Training vs Inference | E1260

  • Jonathan (CEO of Grok) explicitly clarifies that the recently announced $1.5 billion figure is revenue, not a funding round, representing approximately 30% of OpenAI's revenue.
  • Grok is positioning itself to handle low-margin, high-volume inference business (approx. 20% margin) to free up NVIDIA to sell high-margin GPUs for training (70-80% margin).
  • Grok's growth trajectory is described as "faster than exponential," prioritizing market toeholds and relevance over immediate profit maximization.
  • Scaling laws are evolving beyond the traditional logarithmic improvement; the integration of synthetic data generated by smarter models allows for geometric improvements in capability rather than asymptotic drop-offs.
  • DeepSeek's recent efficiency gains are attributed to specific algorithmic improvements (e.g., distilling answers) rather than a fundamental break in the need for compute or data volume.
  • The primary bottleneck for scaling is identified as compute, with data and algorithms viewed as "soft bottlenecks" that can be overcome with sufficient computational power.
  • Grok has expanded its chip count from 640 in early 2024 to over 40,000 by year-end, with a target of over 2 million chips for the current year.
  • Grok's architecture utilizes in-chip memory (LPU) rather than external HBM, improving energy efficiency by roughly 3x and enabling unlimited scalability in capacity compared to GPU limitations.
  • Grok's cost to run inference is more than 5x lower than NVIDIA's GPUs; a single GPU's OpEx (energy/data center) alone equals Grok's total CapEx plus OpEx for equivalent output.
  • Grok secured a deal with a Saudi entity where the partner provides all CapEx, with Grok recouping costs via revenue share once an IRR threshold is met.
  • The company faces a predicted oversupply of data center power in 3-4 years, as current hype-driven gigawatt requests (60x demand) will lead to a market correction once the 18-24 month hardware doubling cycle realizes demand hasn't materialized.
  • Grok's operational model maintains a small team of 300 employees despite massive scaling, relying on automation and "sublinear" employee growth to solve "problem units."
  • Grok explicitly refuses to train its own proprietary foundation models or log user data to ensure customer trust and avoid competing with its partners.
  • China's AI strategy is viewed as constrained by censorship and privacy concerns (fear of "Jack Ma" style repercussions), potentially stifling innovation compared to Western permissiveness.
  • Europe's lag in AI is attributed to a lack of risk-taking culture and regulatory friction; the speaker proposes creating a Special Economic Zone for 1 million "risk-on" entrepreneurs with deregulated labor markets.
  • The speaker warns against "financial diabetes" in society, where AI-induced abundance could reduce human agency and the drive to solve problems if decision-making is outsourced.
  • Future AI advancement stages are defined as: solving hallucinations, enabling agentic sub-goals, unlocking invention (non-probabilistic creation), and finally proxy decision-making.
  • Grok's internal alignment metric is a physical challenge coin engraved with "25 million tokens per second," used to vet every decision against this single growth objective.
  • The speaker believes aging could be significantly slowed or stopped within 10 years, citing the sudden efficacy of GLP-1 drugs like Mounjaro as a precedent for rapid biological breakthroughs.
  • Grok's philosophy on hiring prioritizes mission-oriented talent over maximum salary to prevent attrition to competitors, accepting a "bidding war" disadvantage to ensure long-term loyalty.
  • The speaker advises founders to build for continuous improvement (better algorithms/data) rather than assuming current scaling laws will plateau, drawing parallels to the pre-smartphone era of internet-only startups.
  • NVIDIA is characterized as a monopsony buyer for HBM, but Grok argues that its LPU architecture bypasses the HBM supply constraint entirely, rendering NVIDIA's memory bottleneck irrelevant for inference.