Interview, Podcast
Mark Zuckerberg — Llama 3, $10B models, Caesar Augustus, & 1 GW datacenters
- Meta projects that AI development will reach a state of surpassing human capabilities across most skills over the next two years, characterizing this progression as an additive process rather than a single threshold event.
- The company plans to release a 70 billion parameter model today featuring leading scores in math and reasoning, followed later in the year by a 405 billion parameter dense model currently training at 85 MMU, with a long-term strategy to scale to larger or specialized architectures.
- Future product roadmaps include multimodal models for smart glasses, AI agents capable of managing individual creator and business communities 24/7, and a shift from simple chatbots to systems executing complex, multi-step tasks without intervention.
- Significant infrastructure challenges are anticipated regarding energy consumption and data center construction, with potential bottlenecks arising from the need to build facilities ranging from 300 megawatts to 1 gigawatt, which have not yet been constructed.
- Meta anticipates a fundamental shift in training methodologies where a significant portion of "training" may eventually rely on generating synthetic data via inference, alongside a transition from hand-engineered tool use to integrated model logic.
- A $100 billion investment is underway to secure strategic independence and prevent reliance on competitors, with revenue generation expected through licensing and revenue-share arrangements for cloud providers reselling Meta models.
- The open-source strategy for Llama models aims to prevent dangerous power concentration and commoditize training costs, though Meta may withhold open-sourcing if a qualitative threshold of capability or danger is crossed, and will retain proprietary rights for specific product code like Instagram.
- Risks to be managed include day-to-day harms such as misinformation, fraud, and election interference, while the "runaway" scenario of overnight intelligence explosion is deemed unlikely due to physical energy and infrastructure constraints.
- Current technical progress indicates the 8 billion parameter model is nearly as powerful as the largest Llama 2 version, while the 70 billion parameter model trained on 15 trillion tokens shows continued learning potential beyond its current cutoff, which was a business rather than technical decision.
- Meta is developing smaller, efficient models (e.g., 1 billion to 500 million parameters) for specific use cases and expects custom silicon to eventually support training, though Llama 4 is not currently planned for this architecture.