Lecture
Complete Statistical Theory of Learning (Vladimir Vapnik) | MIT Deep Learning Series
- Statistical learning theory is expected to evolve toward understanding the nature of intelligence rather than merely imitating it, requiring a complete theory that combines data usage with intelligent principles.
- Introducing invariants or smart predicates is predicted to significantly improve performance, with specific claims of reducing error rates from 73% to 7% in diabetes datasets and from 3.1% to 2.9% in digit recognition tasks.
- Utilizing specific invariants like Lie derivatives for digit recognition is expected to enable the creation of digit clones, thereby reducing the dependency on large training datasets.
- It is anticipated that a small set of abstract predicates, estimated at a dozen or a couple of dozen, will be sufficient to describe 2D images and achieve intelligence in recognition tasks.
- Researchers are challenged to match the 0.5% error rate of current deep neural networks while utilizing only 1% of the training data (6,000 observations instead of 60,000) by inventing appropriate smart predicates.
- Increasing the number of predicates is expected to decrease overfitting by reducing the size of the admissible set of functions.
- Understanding specific predicates, such as symmetry, is identified as the key to solving image recognition problems, contrasting with the approach of brute-force data scaling.
- While natural language processing presents greater difficulty than image recognition, the same approach of identifying a limited set of abstract ideas is expected to apply to that domain as well.
- The 31 predicates previously described by Vladimir Propp for folklore serve as an analogous model for finding a small set of predicates for 2D images that reflect the essence of intelligence.
- Machine learning is expected to adopt the method used by physicists to resolve contradictions where invariants fail, utilizing this process to improve approximation and discover underlying laws.