newsfilter.io
Interview

AI in Pharmaceutical R&D with Kim Branson

  • Kim Branson, SVP and Global Head of AI/ML at GSK, joined the company five years ago during a strategic "turning inside out" phase, having previously worked at Vertex and startups.
  • Branson's background combines structural biology (molecular biology, bacterial pathogenesis, X-ray crystallography) with computation, a convergence that began with his PhD work on the computational design of Relenza (zanamivir) in the late 1990s.
  • He identifies the primary challenge in large pharma as organizational culture: overcoming the "innovator's dilemma" by balancing legacy processes with the urgency and pace of startups to prevent AI initiatives from being treated as peripheral experiments.
  • GSK's AI strategy focuses on creating a "sheltered bubble" for rapid capability building before integrating AI into the broader, geographically distributed workflow, aiming to transform the company into an "unstoppable juggernaut" of data and capital.

Current AI Applications in GSK's Drug Discovery Pipeline

  • Target Identification: AI models analyze large genetic databases (GWAS) and clinical imaging to derive "continuous traits" (e.g., liver scarring scores) rather than relying on manual scoring, identifying significant genetic drivers of disease.
  • Functional Interpretation: Machine learning predicts the biological directionality of genetic variants (e.g., increased vs. decreased protein expression) and identifies relevant cell types, accelerating the move from genetic association to mechanistic understanding.
  • Active Learning Systems: GSK employs active learning loops where models generate hypotheses based on literature and genetics, which are then validated through automated biological experiments (e.g., CRISPR/TALON gene perturbations), reducing screen times by a factor of 20 compared to random screening.
  • Computational Pathology: AI analyzes tissue slides to quantify protein expression levels and identify specific cell types with higher precision than human pathologists, freeing experts for higher-order analysis.
  • Clinical Trial Optimization: In a highly instrumented Phase 2 Hepatitis B trial, AI analysis of proteomic data identified a specific patient subset that achieved functional cure upon lowering viral surface antigen levels, enabling targeted treatment strategies.

Strategic Insights for Startups and the Broader Ecosystem

  • Data Moats: Branson emphasizes that unique, proprietary, and generated data is more valuable than algorithmic sophistication; "more data with a simple algorithm" (like Random Forest) often outperforms complex models lacking high-quality data inputs.
  • Validation Benchmarks: Startups should target a 10% improvement over baseline metrics (e.g., AUC or precision-recall curves) rather than statistically significant but negligible 0.1% gains, ensuring the technology delivers meaningful clinical utility.
  • Implementation Friction: Successful commercialization requires robust engineering support, including measures for model drift, on-premise deployment options, and clear confidence intervals for predictions, not just point estimates.
  • Iterative Development: Companies should define the minimum viable performance criteria required for a client to adopt the tool before building the solution, preventing "feature creep" and ensuring market fit.

Future Outlook and Rate-Limiting Factors

  • The Rate-Limiting Step: The industry's primary bottleneck is not data volume, but "data with outcomes" (longitudinal clinical data linking specific interventions to patient results), which remains the rarest and most valuable asset.
  • Predictions for 5 Years:
    • Computational Biomarkers: Every new drug launch will be accompanied by software that predicts baseline patient status and response trajectories more accurately.
    • Mechanistic Integration: AI will increasingly collide with mechanistic biology, utilizing structured priors (e.g., organ system interactions) rather than relying solely on gene expression data.
    • Inference Tables: GSK is building large-scale perturbational data sets to serve as inference tables, allowing models to predict experimental outcomes without conducting the physical experiment.
    • Observational Cohorts: A shift toward large-scale, public-private observational cohorts with deep, longitudinal sampling (genomics + biosamples) to reduce disease heterogeneity and understand immune system dynamics over time.
  • Open Science: GSK actively contributes to the field via public challenges (e.g., Gene Disco) to advance the state of the art, though core competitive clinical outcome data remains proprietary.