Conference Presentation, Lecture
Atul Butte | Driving Public Health: A Trillion Points of Data
- The exponential growth of biomedical data is outpacing traditional scientific publication models; researchers now face pressure to make high-throughput molecular data (genomics, proteomics) publicly available to allow for data reuse and verification.
- By August 2012, the public repository of genomic chips reached 1 million samples, with the dataset doubling every 2–3 years and approaching 2 million publicly available samples.
- The number of biomedical databases is itself growing exponentially, creating an "exponent of an exponent" effect that fundamentally changes data accessibility.
- Open biomedical data serves as a unique global export; unlike physical resources, digital data can be shared infinitely to allow diverse global actors (from Bakersfield to Bangladesh) to develop diagnostics and therapeutics.
- The Precision Medicine Initiative, announced one year prior to the speech, emphasizes that precision medicine is not a substitute for universal basic healthcare but rather a complement that requires foundational access for all.
- Current disease classification systems, specifically the ICD (International Classification of Diseases), rely on outdated nosology and fail to capture molecular heterogeneity, often grouping distinct biological conditions under single codes.
- ICD-10 represents only a starting point; the field requires a shift from code-based classification to molecular-based classification to enable true precision medicine.
- Researchers are integrating public molecular data to predict "drug repurposing" opportunities, matching existing drugs to new disease indications computationally before clinical testing.
- A specific case study involved the antidepressant imipramine, which was computationally predicted to treat small cell lung cancer (5% survival rate); this prediction led to a clinical trial authorized within 15 months.
- The imipramine clinical trial was funded at approximately $50,000 via seed grants, demonstrating that low-cost "garage biotech" efforts can advance patient treatments compared to traditional multi-million dollar pharma timelines.
- UCSF has launched a system-wide initiative to consolidate electronic medical records (EMR) from all five University of California health systems (UCLA, UCSF, Irvine, Davis, San Diego) into a single data warehouse.
- This unified system covers 14.1 million patients, representing approximately 4% of the U.S. population, with the goal of achieving full interoperability across different EMR vendors (EPIC, Allscripts, Cerner).
- The UC Health data warehouse aims to construct dynamic "maps" of patient disease progression and mortality to track transitions between conditions over time, such as the trajectory from heart attack to septicemia.
- These maps are intended to define a new standard for accountable care organizations, shifting focus from individual encounters to continuous, population-level monitoring of all 14 million patients.
- A prototype visualization of these maps displays patient cohorts (represented by dots) moving through disease states, with plans to expand this to a real-time wall of monitors tracking health outcomes and future trajectories for the entire patient population.
- A dedicated clinical data warehouse team of 40 personnel across the UC system has been established to manage the engineering and interoperability challenges of this consolidation.
- Funding and support for these initiatives are provided by the NIH, March of Dimes, JDRF, and HP, alongside internal institutional recruitment and administrative support.