Fireside Chat, Interview
Sarah Tavel: Will Foundation Models Be Commoditised? | E1149
Investment Thesis & Market Structure
- Benchmark focuses exclusively on the application layer, arguing that the vast majority of economic value will be captured here rather than at the infrastructure level.
- The infrastructure layer (foundation models) is predicted to become an oligopoly due to escalating compute costs and power constraints, making each successive model more expensive to train.
- Frontier models will predominantly be closed-source; if a company requires the absolute state-of-the-art model, it will likely be behind a paywall, though open-source may suffice for non-frontier use cases.
- Sarah Tavel notes that while infrastructure costs to train models will rise, the cost to the end-customer for inference should continue to decrease due to competition.
AI as a "Sustaining" vs. "Disruptive" Force
- AI is currently acting as a "sustaining technology" for incumbents (e.g., Notion, Adobe) who can integrate API-based capabilities into existing workflows, leveraging their massive distribution to stifle startups.
- Disruptive opportunities exist only when startups shift from selling "productivity tools" (seats) to "selling the work" (outcomes), bypassing the need for employee training and changing the pricing model.
- Startups must move beyond "wrappers" around LLMs; they must own a larger percentage of the workflow value to defend against incumbents who can simply add features to their existing products.
- The market opportunity size for AI startups selling work is theoretically 10x–50x larger than traditional SaaS because they are selling against the full cost of headcount rather than marginal productivity gains.
Evaluation Criteria & Investment Discipline
- The "Why Now" factor is the single most critical metric; a company must have a sustaining technological or cultural current to drive momentum, otherwise it faces a "paddle against a tide" dynamic.
- Benchmark rejects the industry trend of raising massive rounds (e.g., Cognition AI) for pre-revenue companies unless there is a conviction that the capital will build a hard moat via GPU training, avoiding "FOMO capital deployment."
- Differentiation in crowded AI sectors (e.g., customer service agents) relies heavily on the founder's competitive energy, urgency, and ability to navigate a "land grab" rather than just product features.
- Benchmark practices a "one or two new commitments per year" strategy, prioritizing deep, equal partnerships over scaling the number of investments to maintain founder focus.
- The firm rarely breaks its pricing model, viewing price as a litmus test for conviction; if a founder offers a lower valuation to make the deal work, Benchmark will walk away.
- Benchmark does not prioritize "pro rata" rights; after the initial investment, they are 100% aligned with the founder's immediate needs, avoiding the dilution conflicts common with traditional multi-stage funds.
Founder Psychology & Board Dynamics
- Benchmark's "product" is the partnership itself, which involves direct GP involvement in recruiting and strategy rather than delegating to internal specialist consultants.
- Partners view their role as "truth-seekers" rather than cheerleaders, pushing founders to solve scaling issues before they become critical.
- Selection bias plays a major role in lost deals; Benchmark often declines founders who prioritize platform support or wish to remain independent from a board, as this misaligns with their active partnership model.
- Sarah Tavel identifies Ethereum as her biggest investment miss, stemming from failing to act on the realization of smart contracts' disruptive potential despite understanding the technology early.
- Tavel admits that her belief in the prevalence of modern anti-Semitism was a significant shift in perspective over the last 12 months, noting a "metastasization" of the issue compared to previous years.
Specific Market Observations
- DeepL is cited as a prime example of a company that built enduring value by owning its own foundation model for language while layering complex workflows (e.g., vocabulary lists) that generic APIs cannot replicate.
- Eleven Labs exemplifies the "modality focus" strategy, where specializing in audio allows for a competitive headstart over generalist models like OpenAI's GPT series, which prioritize their core text foundation.
- Bereal is used as a cautionary tale where a strong "Why Now" (authenticity) failed to capture enough consumer minutes against the "black hole" of algorithmic short-form video.
- Chainalysis is highlighted as Benchmark's longest-running board relationship, serving as a case study in long-term founder growth and the compounding value of early, committed partnership.