newsfilter.io
Interview, Fireside Chat

Daphne Koller: Biomedicine and Machine Learning | Lex Fridman Podcast #93

  • Daphne Kohler is a Stanford computer science professor, co-founder of Coursera, and CEO of In-Citro, a company applying machine learning to biomedicine.
  • Her transition from education to health began in 2016, driven by a desire to tackle the development of machine learning applications for human health.
  • Kohler expresses skepticism regarding "curing all major diseases," noting that damage is often done before discovery and regeneration is a "very challenging problem."
  • She estimates current understanding of disease mechanisms varies significantly: some areas may be 70–80% understood, while major conditions like Alzheimer's, schizophrenia, and autism are "closer to zero."
  • Alzheimer's and schizophrenia are not singular diseases but likely "heterogeneous collections of mechanisms" that manifest similarly clinically but differ biologically.
  • Kohler distinguishes between increasing "health span" (living healthily to ~120 years) and immortality, viewing the former as a more realistic and worthy societal goal.
  • Aging processes like DNA damage, protein misfolding, and cellular "wear and tear" contribute to disease, creating a significant overlap between aging and pathology research.
  • Historically, machine learning in health has been limited by a lack of large-scale, high-quality datasets.
  • In-Citro's strategy reverses the traditional pipeline: using biological and chemical tools to generate the specific data required for machine learning, rather than applying ML to existing data as a byproduct.
  • Kohler cites her father's death from an autoimmune lung disease as a catalyst for her focus on drug discovery, noting the lack of viable treatments 12 years ago compared to today's options.
  • Traditional animal models (e.g., mice) often fail because the induced disease mechanisms do not match human pathology; mice do not naturally develop Alzheimer's or schizophrenia.
  • In-Citro utilizes "disease-in-a-dish" models, using induced pluripotent stem cells (iPSCs) derived from human skin to create patient-specific cell types (e.g., neurons, cardiomyocytes).
  • The "Yamanaka factor" allows for the reversal of skin cells to pluripotent stem cells, a process now "almost industrialized" by contract research organizations.
  • Current iPSC libraries number between 5,000 and 10,000 globally, which Kohler notes is not yet sufficient for massive population-scale studies but adequate for perturbation experiments.
  • CRISPR gene editing enables the creation of isogenic pairs (healthy vs. mutated cells from the same donor) to isolate the specific effects of genetic mutations.
  • "Polygenic risk scores" can quantify disease risk variations between individuals by a factor of 10–12, with cellular models providing closer signals to clinical outcomes than genetics alone.
  • Data generation for these models utilizes single-cell RNA sequencing and super-resolution microscopy to turn "squishy" biological tissue into digital data.
  • Machine learning is applied to these datasets to identify molecular disease subtypes, discover novel gene pathways, and screen for interventions that revert diseased cells to a healthy state.
  • The approach is most promising for diseases with a strong genetic basis, contained cell types, and reproducible cellular phenotypes.
  • Emerging technologies like "organoids" (mini-organs like cerebral or liver models) and multi-organ system connections are expanding the scope of tractable diseases.
  • Kohler co-founded Coursera in 2011 after observing 100,000+ enrollments in Stanford's first MOOCs without a marketing campaign.
  • Key pedagogical findings from Coursera include: short video modules (5–7 minutes) outperform hour-long lectures; compressed content is effective; and micro-quizzes improve engagement.
  • MOOCs are unlikely to replace face-to-face education but are becoming essential for continuing education and rapid skill updates in a changing job market.
  • She advises aspiring ML practitioners to master foundational math and statistics before applying models to avoid "turning the crank" on incorrect architectures.
  • Kohler identifies "end-to-end training" and "transfer learning" (learning reusable representations) as the most beautiful and surprising concepts in deep learning.
  • She remains skeptical of current "universal" neural networks, noting that architecture still requires domain-specific insight (e.g., convolutional nets for images vs. language).
  • A critical limitation of current ML is poor calibration of uncertainty; models often express high confidence in incorrect predictions, posing risks in medical diagnosis and autonomous driving.
  • Solutions for uncertainty include Bayesian deep learning, Gaussian processes, and ensemble methods, though this remains an open area of research.
  • Kohler argues that fears of AGI "taking over" are premature given current systems lack the generalization and versatility of a human toddler.
  • She warns that "dumber" systems pose immediate risks due to complexity, unpredictability, and the potential for dangerous feedback loops in critical infrastructure.
  • Security threats include adversarial attacks, data misuse (e.g., face recognition), and the dual-use nature of technologies like CRISPR, which could create bioweapons.
  • Kohler observes that while technology can be abused, global trends show decreasing violence and increasing human rights, suggesting a net positive trajectory.
  • Her definition of life's meaning is to "make a dent in the universe," ensuring the world is better upon her death than upon her birth.
  • She emphasizes that privilege carries a burden to use one's life to benefit humanity, particularly for those born into educated, supportive families.