newsfilter.io
Interview, Fireside Chat, Panel

AI Czar David Sacks Explains the DeepSeek Freak Out

  • The release of DeepSeek's R1 model triggered a global news story and a trillion-dollar decline in market cap, driven primarily by two unique factors:

    • The model is from a Chinese company, intensifying the geopolitical "China vs. US" AI competition narrative.
    • The model is open-source, reigniting the "open-source vs. closed-source" debate and angering users of closed providers like OpenAI.
    • These elements combined fueled attention on platforms like TikTok, where international audiences viewed the release as a check against US dominance.
  • DeepSeek R1 is a "reasoning model" that utilizes reinforcement learning and "chain of thought" processing to break complex problems into sub-tasks, unlike standard base LLMs that provide direct answers.

    • OpenAI released the first reasoning model (O1) approximately four months prior.
    • DeepSeek R1 is the second publicly released full version of this technology, though Google and Anthropic have prototypes or private betas in development.
    • The timing compressed global industry expectations regarding China's lead in AI, shifting the perceived gap from 6–12 months to 3–6 months.
  • The widely circulated claim that DeepSeek trained the model for only $6 million is largely disputed as an apples-to-oranges comparison.

    • The $6 million figure reportedly covers only the final training run, excluding prior R&D, experiments, and compute infrastructure costs.
    • In contrast, comparable final training runs by US companies (e.g., Anthropic, OpenAI) cost in the tens of millions.
    • When comparing "soup-to-nuts" fully loaded costs including hardware, DeepSeek's estimated cluster of ~50,000 NVIDIA Hopper GPUs (including H100s, H800s, and H20s) represents an asset value exceeding $1 billion.
  • DeepSeek's technical achievements are attributed to severe resource constraints necessitating innovation in algorithms and infrastructure.

    • The company abandoned the industry-standard PPO algorithm for a custom variant (likely GRPO) that is more memory-efficient and high-performing.
    • The team bypassed NVIDIA's proprietary CUDA stack, writing directly to bare metal using PTX to gain greater control and reduce costs.
    • These innovations suggest that financial constraints can catalyze technical breakthroughs that unlimited capital might not provoke.
  • The event signals a potential shift in the AI value chain where model commoditization could make the underlying technology a "moat" rather than the primary source of value.

    • Investors like Palmer Luckey and Brad Gerstner argue that the low cost and speed of inference shift value creation upstream or to the application/user layer.
    • Historical parallels to the electricity industry suggest that while power generation companies may not capture the most value, the rest of the economy accrues the majority of benefits through widespread adoption.