Interview, Fireside Chat, Conference Presentation
Why Human Data is Key to AI: Alexandr Wang from Scale AI
a16zAlexandr Wang, David George, Sarah Wing, Alex Wang, Max Wiethe, Dan Morehead, Mike Greenleaf, Alex Rosenberg
- Scale AI intends to produce frontier-level data in partnership with large labs and aims to enable enterprises and governments to utilize proprietary data for AI development, a process described as a major human project expected to evolve into a large-scale experiment akin to "the internet on steroids."
- The industry anticipates closing the current phase of language model development within the next three years, driven by a shift toward research-heavy divergence, alternating cycles of execution and innovation, and a reliance on breakthroughs in data production now that accessible public data is exhausted.
- To achieve data abundance, the sector will focus on increasing complexity, generating high-quality agent data, investing in synthetic and hybrid human-in-the-loop data, establishing "data foundries," and adopting more scientific measurement methods to identify specific model gaps.
- Regulatory perspectives on large tech data advantages in Europe remain undetermined, while large labs face a critical risk-reward dynamic where missing AI leadership poses an existential threat to their business models.
- Large tech companies are expected to recoup capital expenditures through efficiency gains in core businesses, such as improved GPU utilization for advertising or driving hardware upgrade cycles, while the surplus from open-source models is forecast to be immense.
- Intelligence is predicted to become a commodity with inference pricing falling by orders of magnitude over two years, likely changing the market structure if open-sourcing continues or performance levels converge, rendering the rental of pure models a mediocre long-term business absent durable breakthroughs.
- Businesses providing underlying infrastructure like Nvidia and cloud providers are expected to remain high-quality due to scale and logistics, while application businesses like ChatGPT will remain strong if they achieve early product-market fit where value exceeds inference costs.
- Major labs will pursue deeper product integrations beyond basic chatbots, though product iteration cycles will remain difficult to predict; OpenAI and Anthropic are expected to need strong application businesses to ensure long-term independence and sustainability.
- The current application layer is in "phase one" characterized by chatbots and partial automation, with near-term benefits focused on cost savings and efficiency gains rather than immediate disruption, as enterprise implementation is a multi-year investment cycle.
- Almost every enterprise holds latent potential to significantly boost stock prices through AI, with AI poised to be the first wave to massively transform products; a race is anticipated between enterprises leveraging internal data and startups accessing data subsets to create distinct products.
- High-performing startups face the risk of "mean regression" when scaling headcount, requiring careful integration of new executives who must understand company rhythms rather than making sweeping changes, while founder-led organizations retain an innovation premium.
- Scale AI commits to hiring based on merit and excellence without demographic quotas, while AGI is defined as AI capable of accomplishing 80% of digital-focused jobs, with an estimated timeline of "four plus years" though algorithmic innovation could accelerate this significantly.