Podcast, Other
The data black hole at the center of AI
Core Thesis: Data Volume Over Sample Efficiency
- Current AI progress is driven primarily by widening and improving data distributions and scaling compute, rather than improvements in training sample efficiency.
- The primary mechanism for generating high-quality data is Reinforcement Learning (RL), which treats the process as synthetic data generation using compute-heavy verification rubrics.
- Success in RL depends on models having a prior probability of anticipating correct solutions, necessitating vast amounts of human expert trajectories for every specific skill domain.
- The resulting industry of human expert labeling and RL environment creation is currently generating billions in revenue with projected decacollar growth.
Characteristics of the Current Training Paradigm
- Data Specificity: Expert data is highly bespoke, requiring domain-specific specialists (e.g., word specialists, legal M&A experts, management consultants) to generate example completions and chain-of-thought explanations.
- Data Volume Disparity: Frontier models are trained on tens to hundreds of trillions of tokens, a volume approximately one million times greater than the ~200 million tokens a human processes in a lifetime.
- Model Architecture Metaphor: Current models are described as "Frankenstein's monsters" composed of billions of carefully constructed graphs sewn together, rather than possessing human-like integrated skill sets.
- Open Source Convergence: Open-source models lag frontier models by approximately four months, a gap closable because data (easily distilled from public APIs) drives progress more than proprietary hyperparameters or architectural tricks.
Analysis of Sample Efficiency Constraints
- Robotics Gap: The inability to replicate human-level sample efficiency prevents AI from learning to operate humanoid robots in hours, a necessary condition for a deca-trillion dollar robotics industry.
- Autonomous Driving: Self-driving models (Waymo, Tesla) utilize data volumes three to four orders of magnitude higher than a human's ~20 hours of driving practice, despite human physical intuition.
- Evolutionary Objection Refuted: The human genome (3 gigabytes) is insufficient to encode the parameters of a neural net; evolution likely optimized hyperparameters/loss functions, while humans still build their "connectome" (weights) from scratch during their lifetimes.
- Scaling Law Limitations: Chinchilla scaling laws indicate that even infinite parameter scaling would only reduce the required data by a factor of 10, failing to bridge the thousands-to-millions-of-times efficiency gap between humans and current models.
- Multimodal Data Rebuttal: Sensory data from sight/hearing accounts for only tens to hundreds of billions of tokens; blind/deaf individuals possess general intelligence with significantly less data, suggesting sensory input is not the primary driver of human sample efficiency.
Economic and Strategic Implications
- White-Collar Automation Strategy: Labs assume common tasks (software engineering, accounting, analysis) are frequent enough to be included in training distributions, making inefficient training viable due to the ability to amortize costs across billions of sessions.
- Job Market Projection: Despite AI automation, the demand for human software engineers is projected to increase by 2027 due to AI acting as a complementary input rather than a pure replacement.
- Future Roadmap: The long-term plan involves first automating AI research itself, with the goal of having automated AI researchers solve the fundamental sample efficiency problem.
- Forward-Looking Uncertainty: It remains an open, complex question whether AI can solve the remaining research bottlenecks given the current "clumsy" understanding of how AI-led AI progress would manifest.
Advertisement Segment: Mercury Command
- Product Functionality: "Command" is an AI tool integrated into the Mercury banking platform designed to automate financial planning and tax preparation.
- Operational Logic: The tool analyzes current balances, upcoming invoices, and six months of transaction history to calculate monthly averages and scheduled payments.
- User Interaction: Users can flag non-platform items (e.g., external contractor payments) via chat; the system generates transfer drafts for approval with links to underlying data for verification.
- Compliance Disclosures: Mercury is a FinTech, not a bank; banking services are provided through Choice Financial Group and Column NA (FDIC members); AI suggestions are not guaranteed.