Presentation, Lecture
Turing Test: Can Machines Think?
- Gary Tan plans to overview Alan Turing's paper, consider objections, and discuss alternatives, predicting the paper will be the most impactful in AI history and the seed for engineering breakthroughs from the 1930s to deep learning.
- Expectations include the Turing paper spawning an immeasurable number of researchers, leading to human-level collective intelligence and making "thinking machine" a non-contradictory phrase by 2000 (50 years post-paper).
- By 2000, a machine with 100 megabytes of storage is predicted to fool 30% of humans in a five-minute conversation test, with human-level AI becoming so commonplace it is taken for granted.
- Learning machines are expected to be a critical and central component for achieving human-level conversational capabilities, though significant work remains in the field.
- Future developments favor end-to-end learning-based approaches for open-domain conversation, with humor identified as one of the hardest aspects of human-level intelligence to achieve.
- Risks and limitations involve the potential for the future of AI to be a journey rather than a destination, with the Total Turing Test potentially being harder to pass than other modalities due to bandwidth constraints.
- The 2014 event where Eugene Guzman fooled 33% of judges is expected to have relied on "smoke and mirrors" and tricks rather than deep conversation, occurring without rigorous, open-domain testing.
- Current challenges include the lack of broad interest from major groups like Google Deep Mind and Facebook AI in the Turing test, the discontinuation of Lobner Prize funding, and the surprise that Amazon limited the Alexa Prize to students only.
- The MENA system with 2.6 billion parameters is a serious open-domain attempt, but its reported results of 79% vs 86% human performance should be taken with a grain of salt due to closed-source methodology and potential PR bias.
- Alternative tests face scrutiny: the Lovelace test is difficult to formalize regarding "surprise," the Lovelace 2.0 test is viewed as more subjective than the Turing test, and the Hutter Prize is considered a good challenge but lacks a poetic "impressed" bar.
- Researchers are expected to take the Alexa Prize and Lobner Prize seriously, as analyzing natural language conversation stands is the real test for human-level intelligence, with no team having passed the Turing test as constructed by the Alexa Prize to date.
- The Abstraction and Reasoning Challenge (ARC) is noted as interesting with a May deadline, though the speaker has not fully internalized it, and semi-autonomous vehicles are expected to remain while human-robot interaction problems require solutions.
- The "Lex Plus AI Podcast" Discord community anticipates contributions of art, code, and ideas, with meetings aiming to host 300-400 people for intimate discussion rather than one-person presentations.
- The club expects to balance accessibility for high school students with utility for expert researchers, prioritizing beautiful and impactful insights over full coverage of paper contents.