newsfilter.io

Eric Jang

Showing 11 of 1 transcripts.

  1. Dwarkesh Patel2h 37m

    What rebuilding AlphaGo teaches us about self-play, RL, and future of LLMs - Eric Jang

    Eric Jang, Ron Minsky, Dan Pontecorvo

    Eric Zhang reconstructs AlphaGo to demonstrate how modern computing, including LLM-assisted coding and efficient neural architectures, reduces training costs from millions to thousands of dollars while solving Go's NP-hard complexity through Monte Carlo Tree Search. The presentation details the evolution from human-supervised data to tabula rasa self-play, highlighting how MCTS provides low-variance supervision that stabilizes value function learning for mid-game states. This framework validates Go as a scalable sandbox for testing automated AI research, offering transferable insights for robotics and drug discovery via verifiable performance loops.