newsfilter.io
Conference Presentation, Keynote

Michael Jordan

  • Big data volumes and velocities are projected to increase in complexity by thousands of times within ten years, creating a critical need for systems that maintain constant run times and error rates despite scaling to billions of models rather than single omnibus ones.
  • Many current personalized business models face statistical failure due to the inability to transfer statistics between individuals with varying data amounts, a challenge requiring a shift from controlling the L2 norm to the L infinity norm to establish a mathematically distinct inferential world.
  • Significant uncertainty remains regarding the existence of current engineering principles to keep performance constant as resources increase, as complexity theory suggests run times may grow proportionally to n, n log n, or n cubed while the industry currently relies on ad-hoc "hacking."
  • The integration of privacy concerns necessitates new mathematical constructs to balance externalities, with the industry estimated to be decades away from having real principles capable of simultaneously addressing personalized service, privacy, and performance.
  • A fundamental disconnect exists between computer science and statistics, described as an "oil and water problem," because core statistical theory lacks direct run time concepts and complexity theory rarely addresses statistical risk.
  • Future theoretical developments will likely require blending computational thinking (modularity, abstraction, robustness) with inferential thinking (sampling patterns) and integrating geometry, information theory, and concurrency control to bridge statistical risk and computational quantities.
  • A new minimax theory has been developed that integrates statistical decision theory with differential privacy, predicting the effective sample size will be calculated by multiplying data points by the square of the alpha parameter and dividing by the problem dimension.
  • Addressing the inability of standard bootstrapping to scale to terabyte-sized data, the "Bag of Little Bootstraps" method uses subsampling from histograms to reduce footprints, achieving error bar quality in a couple of hundred seconds compared to approximately 15,000 seconds for standard methods.
  • Frequentist approaches are deemed necessary for real-life decision-making contexts like medical testing where uncertainty quantification is critical, whereas Bayesian inference remains sensitive to unknown priors and tail behavior in high dimensions.
  • Current software limitations require a transition from descriptive statistics that reference existing databases to inferential thinking that applies models to new, unknown populations, with the ultimate goal of proving high-probability closeness between an estimator and its private version.