newsfilter.io
Lecture, Conference Presentation

Privacy Preserving AI (Andrew Trask) | MIT Deep Learning Series

  • Combining remote execution, search, sampling, and differential privacy is expected to enable pip-installable access to previously inaccessible datasets, allowing questions to be answered using unseen data once the underlying theory is paired with continued tooling efforts.
  • The outlook predicts that several orders of magnitude more data from enterprise and government warehouses will become available relatively quickly for scientific progress without increasing supply scarcity, replicating the progress driven by new big data sets in the past.
  • Infrastructure development for personal privacy budgeting is forecasted to arrive in waves, beginning with enterprise adoption driven by commercial utility rather than privacy narratives, before evolving into global accounting mechanisms resembling IRS-like data banks.
  • The speaker anticipates the emergence of one thousand startups based on a repeatable business model creating gateways between individual data holders and the global market, contingent on building secure infrastructure capable of handling information on major world problems.
  • Regulatory adoption is indicated by the US Census utilizing differential privacy for 2020 data, with the first pilot programs for the OpenMind community and related technologies expected to roll out within the current year.
  • End-to-end encrypted services are envisioned to eventually permit medical diagnoses without revealing records, facilitate holistic recommendation systems using private metrics like sleep quality, and enable unbiased model adjustments by provisioning privacy budgets to measure bias without raw data access.
  • Significant technical barriers currently exist, including a 13x slowdown in deep learning prediction using secure multi-party computation and high compute and network overhead for encrypted services, though hardware optimization similar to NVIDIA's impact on standard deep learning is expected to mitigate these issues.
  • Existing business models involving the purchase, de-anonymization, and resale of anonymized data sets are identified as a risk, while federated learning alone is noted as insufficient for security without the addition of differential privacy to prevent data memorization.
  • The realization of single-use accountability for current accountability systems and the ability to perform data science without direct data access are presented as future outcomes dependent on the successful deployment of these new tools.
  • Full realization of personal privacy budgeting infrastructure is predicted to be a long-term endeavor, with the speaker noting that investment flows are currently hindered by the lack of a direct path to value in personal privacy infrastructure.