Yoshua Bengio
Showing 1–15 of 17 transcripts.
- 80,000 Hours2h 35m
Godfather of AI: How To Make Safe Superintelligent AI – Yoshua Bengio
Yoshua Bengio proposes "Scientist AI," a new paradigm developed by the startup LawZero that trains models to approximate a Bayesian posterior, effectively distinguishing verified truth from human speech acts to ensure honesty by design. By replacing standard reinforcement learning with a loss function that penalizes deviations from verified facts like mathematical proofs and code outputs, the system aims to eliminate deceptive instrumental goals while potentially increasing capability through better causal reasoning. To validate this approach, the organization has raised $35 million to deploy non-agentic safety guardrails within months, advocating for international coalitions to fund the technology and prevent a global race to the bottom on AI safety.
- The Diary Of A CEO1h 40m
Godfather of AI: We Have 2 Years Before Everything Changes!
YOSHUA BENGIO, Steven, Stephen Bartlett
Deep learning pioneer Professor Yoshua Bengio has shifted from technical research to public advocacy, warning that current AI trajectories risk catastrophic harm through misalignment, physical threats, and the democratization of destructive knowledge. Driven by a precautionary principle rooted in concern for future generations, Bengio argues that market incentives and geopolitical races are accelerating these dangers while regulatory measures remain inadequate. To counter these threats, he founded the Law Zero nonprofit to develop "safe by construction" systems and is calling for international treaties, mandatory liability insurance, and shifted public opinion to enforce stricter global safety standards.
- Lex Fridman1h 46m
Michael I. Jordan: Machine Learning, Recommender Systems, and Future of AI | Lex Fridman Podcast #74
Michael I. Jordan, Lex Fridman, Andrew Ng, Zoubin Ghahramani, Ben Taskar, Yoshua Bengio, Yann LeCun
Michael I. Jordan reframes the current state of artificial intelligence not as the engineering of human-like cognition, but as a nascent discipline focused on building large-scale decision systems, while explicitly rejecting premature claims of deep neurological understanding or full brain-computer integration. He distinguishes his approach from pure prediction by prioritizing decision-making under uncertainty and advocates for a shift from ad-based surveillance economies to direct producer-consumer markets that utilize game theory to align incentives with societal health. Jordan concludes that advancing this field requires a blend of rigorous mathematical frameworks, such as empirical Bayesian methods, and broad humanistic education to cultivate the empathy and collaboration necessary for solving unsolved challenges like natural language understanding.
- Lex Fridman1h 28m
Deep Learning State of the Art (2020)
Pamela McCordick, Alan Turing, Frank Rosenblatt, Yann LeCun, Geoffrey Hinton, Yoshua Bengio, Walter Pitts, Warren McCulloch, Alexei Evaknenko, V.G. Lapa, John Hopfield, Juergen Schmidhuber, Rodney Brooks, Sebastian Reuter, Jacob, Noah Brown, Chris Ferguson, Darren Elias, Jeremy Howard, Ian Goodfellow, Aaron Corville, Andrew Trask, Francois Chollet, David Silver, Robbie Allen, Victor Flevin, Ilias Esquiver, Peter Singer, George Washington, Stalin
This presentation traces the evolution of artificial intelligence from Alan Turing's foundational predictions to 2019's deep learning dominance by LeCun, Hinton, and Bengio, while analyzing recent paradigm shifts in reinforcement learning and autonomous vehicle strategies. The speaker highlights 2020's framework convergence between TensorFlow and PyTorch, details the limitations of current transformer-based models regarding common sense reasoning, and outlines critical research priorities in ethics and long-term safety. Ultimately, the discourse frames the greatest existential risk not as rogue AI, but as human utilization of these tools for control and warfare, urging a democratization of the technology to ensure ethical stewardship.
- Lex Fridman42 min
Yoshua Bengio: Deep Learning | Lex Fridman Podcast #4
This event synthesizes current limitations in artificial neural networks, highlighting the need to shift from passive observation to active agent learning and the development of disentangled representations for better causal reasoning. It proposes that future progress depends on integrating unsupervised semantic understanding with supervised labels and leveraging machine teaching strategies to mimic human attention mechanisms. Furthermore, the discussion reframes AI safety priorities away from fictional existential threats toward immediate societal challenges like algorithmic bias, autonomous weapons, and the ethical alignment of systems through robust world models.
- Lex Fridman1h 12m
Foundations and Challenges of Deep Learning (Yoshua Bengio)
Yoshua Bengio, Hugo Larochelle, Andrej Karpathy, Richard Socher, Sherry Moore, Ruslan Salakhutdinov, Andrew Ng, John Schulman, Pascal Lamblin, Adam Coates, Alex Wiltschko, Quoc Le, Shubho Sengupta
Yoshua Bengio outlines five essential ingredients for human-level machine learning, emphasizing that deep neural networks overcome the curse of dimensionality through parallel and sequential composition to efficiently represent complex functions. He contrasts current high-dimensional optimization landscapes, which are dominated by saddle points rather than local minima, against historical theories while highlighting unsupervised learning as a critical mechanism for developing generalizable world models. The presentation concludes by addressing future challenges in training long-term dependencies and integrating neuroscience-inspired alternatives to backpropagation, alongside administrative notes regarding an upcoming textbook by Bengio, Ian Goodfellow, and Aaron Courville.
- Lex Fridman1h 25m
Foundations of Unsupervised Deep Learning (Ruslan Salakhutdinov, CMU)
Ruslan Salakhutdinov, Hugo Larochelle, Andrej Karpathy, Richard Socher, Sherry Moore, Andrew Ng, John Schulman, Pascal Lamblin, Adam Coates, Alex Wiltschko, Quoc Le, Yoshua Bengio, Shubho Sengupta
This presentation details the evolution of unsupervised learning from sparse coding and autoencoders to complex probabilistic frameworks like Restricted Boltzmann Machines, Variational Autoencoders, and Generative Adversarial Networks. It highlights how these non-probabilistic and probabilistic models overcome the scarcity of labeled data by learning hierarchical representations, with GANs notably producing sharper images than VAEs by avoiding explicit density estimation. The discussion further illustrates practical applications ranging from multimodal image-text modeling and semantic vector arithmetic to one-shot learning capabilities.
- Lex Fridman57 min
Torch Tutorial (Alex Wiltschko, Twitter)
Alex Wiltschko, Hugo Larochelle, Andrej Karpathy, Richard Socher, Sherry Moore, Ruslan Salakhutdinov, Andrew Ng, John Schulman, Pascal Lamblin, Adam Coates, Quoc Le, Yoshua Bengio, Shubho Sengupta
This presentation details the practical implementation and theoretical foundations of the Torch deep learning framework using the Lua language, developed in collaboration with experts from Facebook, Google, and Twitter. The speaker explains how Torch leverages LuaJIT for high-performance embedded deployment while utilizing its dynamic Autograd system to support flexible control flow and custom gradients without the overhead of static computation graphs. Case studies from Twitter demonstrate the framework's transition from a research tool for cutting-edge models like GANs to a production environment for serving media, highlighting its efficiency in both training via reverse-mode differentiation and inference through lightweight C++ integration.
- Lex Fridman1h 21m
Sequence to Sequence Deep Learning (Quoc Le, Google)
Quoc Le, Hugo Larochelle, Andrej Karpathy, Richard Socher, Sherry Moore, Ruslan Salakhutdinov, Andrew Ng, John Schulman, Pascal Lamblin, Adam Coates, Alex Wiltschko, Yoshua Bengio, Shubho Sengupta
This presentation details the evolution of sequence-to-sequence learning for automating email responses, transitioning from bag-of-words models to Recurrent Neural Networks and advanced attention mechanisms. Key technical advancements include the encoder-decoder architecture with beam search decoding, personalized user embeddings, and gated units like LSTMs to manage long-term dependencies and vocabulary limitations. The discussion concludes by highlighting real-world applications in machine translation and conversational AI, alongside future research directions in unsupervised learning and global sequence optimization.
- Lex Fridman1h 29m
Deep Learning for Natural Language Processing (Richard Socher, Salesforce)
Richard Socher, Hugo Larochelle, Andrej Karpathy, Sherry Moore, Ruslan Salakhutdinov, Andrew Ng, John Schulman, Pascal Lamblin, Adam Coates, Alex Wiltschko, Quoc Le, Yoshua Bengio, Shubho Sengupta
This presentation outlines the evolution of Natural Language Processing from hierarchical linguistic analysis to deep learning architectures like Word2Vec, GRUs, and Dynamic Memory Networks that handle sequence modeling and visual question answering. Key researchers discussed how continuous vector representations and gated units address ambiguity and long-range dependencies, achieving state-of-the-art performance on benchmarks such as Facebook's bAbI dataset and reducing language modeling perplexity to 70. Despite these advances, the discussion highlights persistent challenges regarding unified joint models, data scarcity in specialized domains, and the need for improved robustness against adversarial inputs and false premises.
- Lex Fridman1h 27m
Deep Reinforcement Learning (John Schulman, OpenAI)
John Schulman, Hugo Larochelle, Andrej Karpathy, Richard Socher, Sherry Moore, Ruslan Salakhutdinov, Andrew Ng, Pascal Lamblin, Adam Coates, Alex Wiltschko, Quoc Le, Yoshua Bengio, Shubho Sengupta
This technical presentation delineates Deep Reinforcement Learning as a sequential decision-making framework that employs neural networks to maximize cumulative rewards through policy gradients and Q-function learning. Key figures in the field, such as those at DeepMind, have leveraged these methods to master complex environments including Atari games, Go, and robotic locomotion by addressing challenges like reward sparsity and non-stationary state dynamics. The discussion further contrasts algorithmic trade-offs between sample efficiency and robustness while outlining future directions like hierarchical structures and model-based approaches to enhance real-world deployment.
- Lex Fridman1h 20m
Nuts and Bolts of Applying Deep Learning (Andrew Ng)
Andrew Ng, Hugo Larochelle, Andrej Karpathy, Richard Socher, Sherry Moore, Ruslan Salakhutdinov, John Schulman, Pascal Lamblin, Adam Coates, Alex Wiltschko, Quoc Le, Yoshua Bengio, Shubho Sengupta, lexfridman, Peter, Andre, Shubo, Sammy
Baidu structures its 1,000-person AI organization around unified data warehouses and integrated ML-HPC teams to drive deep learning performance that scales linearly with data volume rather than traditional algorithms. The presentation outlines critical diagnostic frameworks for bias and variance, emphasizing human-level error as a benchmark for defining theoretical limits and guiding the shift toward end-to-end learning in data-rich perception tasks. Finally, the discussion establishes practical heuristics for product automation and career development, advocating for synthetic data engineering and the rigorous "dirty work" of replicating research papers to master the field.
- Lex Fridman1h 2m
TensorFlow Tutorial (Sherry Moore, Google Brain)
Sherry Moore, Hugo Larochelle, Andrej Karpathy, Richard Socher, Ruslan Salakhutdinov, Andrew Ng, John Schulman, Pascal Lamblin, Adam Coates, Alex Wiltschko, Quoc Le, Yoshua Bengio, Shubho Sengupta, lexfridman, Zach, Pichin Lo
Google Brain's Sherry Moore presented a tutorial on transitioning from research to production using the TensorFlow framework, highlighting its open-source architecture that supports diverse applications like image recognition, voice processing, and deep learning. The session detailed core concepts such as data flow graphs, placeholders, and session execution while guiding attendees through hands-on labs for linear regression and MNIST digit classification. Moore also outlined the platform's extensive portability across mobile and cloud devices and invited community contributions to further develop the library's modular design.
- Lex Fridman1h 25m
Deep Learning for Computer Vision (Andrej Karpathy, OpenAI)
Andrej Karpathy, Hugo Larochelle, Richard Socher, Sherry Moore, Ruslan Salakhutdinov, Andrew Ng, John Schulman, Pascal Lamblin, Adam Coates, Alex Wiltschko, Quoc Le, Yoshua Bengio, Shubho Sengupta, lexfridman
This presentation traces the evolution of convolutional neural networks from 1960s neuroscience foundations to the 2012 AlexNet breakthrough, highlighting how deep learning displaced traditional feature extraction by achieving near-human accuracy on the ImageNet dataset. Key architectural innovations, such as residual skip connections in ResNets and efficient Inception modules, enabled the training of deeper, wider models that serve as generic feature extractors for diverse tasks ranging from object detection to medical imaging. The discussion concludes with practical deployment strategies emphasizing the use of pre-trained models and GPU-accelerated infrastructure to overcome computational bottlenecks in both cloud and edge environments.
- Lex Fridman1h 3m
Theano Tutorial (Pascal Lamblin, MILA)
Pascal Lamblin, Hugo Larochelle, Andrej Karpathy, Richard Socher, Sherry Moore, Ruslan Salakhutdinov, Andrew Ng, John Schulman, Adam Coates, Alex Wiltschko, Quoc Le, Yoshua Bengio, Shubho Sengupta
This presentation details Theano, an eight-year-old symbolic expression compiler that enables high-performance deep learning by automatically differentiating mathematical graphs and compiling them into optimized C++ or CUDA code. The session demonstrates the framework's core capabilities, including graph manipulation for neural network backpropagation, GPU acceleration via shared variables, and sequence modeling through the `scan` operator, while showcasing practical implementations of logistic regression, convolutional networks, and LSTMs on datasets like MNIST. Addressing deployment challenges inherent in its tight Python integration, the discussion concludes by highlighting Docker containers as the standard solution for distributing models and outlines a roadmap for enhanced 3D convolution and cuDNN support.