newsfilter.io
Interview, Fireside Chat

Chai Discovery's Bitter Lesson: Drug Design Is Another Scaling Problem

  • Company Mission & Approach

    • Chai Discovery aims to industrialize drug discovery by replacing serendipitous "trial and error" with a computer-aided design framework for molecules, similar to engineering software code.
    • The founders reject the "bespoke problem" view of biology, treating amino acid sequences as prompts for foundational AI models that can be scaled and generalized.
    • Decision: The company adopted a business model focused on building infrastructure and providing tools to pharmaceutical partners (e.g., Eli Lilly, Novartis, Pfizer) rather than developing a full-stack internal drug pipeline.
    • Reasoning: This "infrastructure" approach allows Chai to allocate the majority of capital toward improving model quality and scaling compute, rather than diverting resources into high-cost clinical trials and drug-specific pipelines.
  • Technical Evolution & Model Capabilities

    • Chai 1: Featured 23 distinct sub-modules, creating complexity that hindered iteration and scaling.
    • Simplification Strategy: Chai pivoted to a simpler architecture, removing unnecessary modules to identify core scaling directions and improve model dynamics.
    • Chai 2 Performance:
      • Achieved a binding success rate of ~15% (150 hits per 1,000 molecules), a massive increase from the industry baseline of 0.1% (1 in 1,000).
      • Enabled "zero-shot" molecule generation, allowing users to specify design principles upfront without needing to screen vast libraries in the lab first.
    • Diffusion Models: The team utilizes diffusion models (rather than Variational Autoencoders) to iteratively "fix" noisy protein structures, allowing the model to learn shortcuts for incremental improvements.
    • Data Sources: Models are trained from scratch on massive sequence databases (potentially exceeding English text token counts) and structural data from the Protein Data Bank (PDB).
    • Feedback Loop: Wet lab results from experiments and partner programs are used to generate new training data, creating a compounding flywheel for model improvement.
  • Scientific Milestones & Hard Targets

    • Antibody Design: Overcame the skepticism that antibody design was too data-scarce, successfully deploying models to design therapeutic antibodies that rival traditional methods.
    • Unlocking Novel Biology: The team believes scaling laws will naturally unlock "undruggable" targets (e.g., specific GPCRs, glycosylation sites) without needing specific engineered modules for each unique target.
    • Developmental Properties: Newer models now bake in manufacturability and drug-like properties at the generation stage, reducing the need for downstream chemical "clean-up."
  • Team Composition & Hurdles

    • Interdisciplinary "Avengers" Squad: The team has evolved from a core of AI researchers to include top-tier antibody engineers (e.g., Andy Young), biologists, product designers (e.g., ex-Stripe, ex-Munz), and GPU infrastructure experts.
    • Hiring Philosophy: Chai hires for the "next milestone" rather than a fixed role; team members often transition from AI research to specific biological applications as the model capabilities mature.
    • Operational Challenges:
      • Scale: Managing large GPU clusters (moving from 128 GPUs to massive scales) introduces complex infrastructure issues like heat management, cluster health, and automatic recovery of long-running training jobs.
      • Rigor: The partner model requires extreme rigor; models must work reliably upon delivery, as partners do not tolerate the "tech debt" common in pure research settings.
  • Industry Outlook & Vision (2035–2100)

    • Future of Drug Development: The industry will shift from "first in class" or "best in class" to "last in class," where AI-generated medicines are the final, optimal answer to specific diseases.
    • Economic Impact: Faster iteration loops (moving from 9 months to 9 weeks or 9 days) will make it economically viable to pursue rare diseases, personalized medicines, and historically undruggable targets like Alzheimer's.
    • Safety & Specificity: Future drugs will be designed with extreme specificity to avoid negative interactions, a level of control achievable only through computational modeling before clinical testing.
    • Competitive Landscape: Chai views the primary competitor as nature's baseline; they aim to clear the bar set by decades of wet-lab optimization (e.g., yeast display) before competing with other AI models.
  • Naming & Culture

    • Name Origin: "Chai" stands for Chemistry + AI, chosen to reflect the company's guiding principle of simplicity and to avoid the complexity of traditional biotech names.
    • Culture:
      • Best Aspect: High mission-driven motivation and a shared clear philosophy on solving hard problems.
      • Worst Aspect: The pain of scaling infrastructure and the pressure of maintaining production-level code quality for enterprise partners, often requiring deliberate "slowing down" in the short term to ensure long-term stability.