newsfilter.io
Interview

AI at the Intersection of Bio | Vijay Pande, Surya Ganguli & Bowen Liu

  • Current State of AI in Drug Design: The field has shifted from questioning if AI can drive drug design to determining how to operationalize it; this transition required decades of evolution.
  • Evolution of Methodologies: Early computational biology (40+ years) relied on physics-based methods (generalizable but computationally expensive) and expert systems (efficient but not generalizable).
  • Deep Learning Revolution: The advent of deep learning allows models to learn raw representations without manual feature engineering, overcoming the limitations of traditional machine learning approaches.
  • Data Scale and Self-Supervision: Massive unlabeled datasets enable self-supervised learning (predicting next tokens), with GPT-4 trained on approximately 5 trillion token sequences.
  • Protein Sequence Data: Evolutionary Scale Modeling (ESM) applies language modeling to 2.8 billion amino acid sequences (roughly 1 trillion tokens), providing a data scale comparable to human language models.
  • Data Sparsity in 3D Structures: Solved protein structures in the Protein Data Bank number only ~200,000, representing significantly less data than genomic sequences.
  • Chemical Space Complexity: The total chemical space is ~10^180, while the accessible drug-like space is ~10^40, necessitating generative AI to explore regions unseen in commercial catalogs (~10^7–10^8 molecules).
  • Generative AI Challenges: Unlike text or images, validating generated scientific concepts is difficult; the primary bottleneck remains the experimental synthesis and testing of AI-generated candidates.
  • Validation Benchmarks: Current generative models often show false positive rates of 20–30% for activity, a regime that is inefficient compared to idealized success rates where 4–5/5 generated candidates are active.
  • Industry Economics: The cost of bringing a new drug to market is ~$2.5 billion over 10–15 years, with a 90% failure rate; the number of drugs approved per billion dollars spent is halved every nine years (inverse Moore's law).
  • Target Coverage: FDA-approved drugs currently target only ~800 of the ~20,000–25,000 human genes, leaving a vast space for exploration.
  • Model Generalization: Physics-based methods often outperform ML models (e.g., AlphaFold) when test data distributions differ significantly from training data, highlighting ML's struggle with out-of-distribution generalization.
  • Optimal Modeling Strategy: The best approach for specific tasks is a specialized model trained on relevant data; a foundation model followed by fine-tuning is the second-best option.
  • Latent Space Theory: Theories suggest biological tasks are combinations of a finite set of underlying skills; finding the correct latent space allows systems to "interpolate" across tasks that appear to be extrapolations (e.g., predicting planetary motion from falling apples).
  • Single-Cell Foundation Models: Models trained on 36 million single-cell RNA-seq expressions create embedding spaces that can hold-out species and map how drugs or diseases shift cellular states.
  • Disentangled Latent Spaces: Variational autoencoders can learn interpretable directions in latent space (e.g., mimicking facial expressions) to design drugs that move biology in specific, desirable directions.
  • Clinical Trial Optimization: AI can improve the 80% failure rate of trials (often due to enrollment issues) by using EMR and biomarker data to reduce patient heterogeneity and select optimal populations.
  • Personalized Medicine: Induced pluripotent stem cell (iPSC) technology enables the creation of patient-specific tissues (e.g., heart tissue) to predict cardiotoxicity and differential drug responses.
  • The "Digital Human" Vision: The long-term goal is a mega-foundation model integrating multi-modal data (proteomics to wearables) to simulate human responses, predict clinical outcomes, and guide personalized treatment.
  • Timeline: Implementation of these integrated AI systems in clinical settings is estimated to take approximately 10 years, though incremental advancements are occurring now.
  • Regulatory Requirements: AI-driven patient selection and trial design will require high interpretability to satisfy FDA regulators regarding decision-making processes.
AI at the Intersection of Bio | Vijay Pande, Surya Ganguli & Bowen Liu — Summary