Latest Interviews
Showing 1–3 of 3 transcripts.
Clear all filters- a16z42 min
Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building
Jack Parker-Holder, Shlomi Fruchter, Anjney Midha, Marco Mascorro, Justine Moore, Erik Torenberg
Google DeepMind has released Genie 3, a research preview that generates interactive, photorealistic 3D worlds in real-time from text prompts to support navigation and control. Built by integrating insights from three internal projects, the model introduces spatial memory for one-minute object persistence and emergent physical reasoning to distinguish it from previous video generation systems. While currently limited to visual simulation without audio, Genie 3 aims to bridge the sim-to-real gap for robotics and agent training by providing diverse, high-fidelity environments free from physical data collection risks.
- a16z1h 18m
Rick Rubin: Vibe Coding is the Punk Rock of Software
Rick Rubin, Marc Andreessen, Ben Horowitz, Anjney Midha, Erik Torenberg
Rick Rubin introduces "vibe coding" as a methodology merging the ancient spiritual principles of the Tao Te Ching with modern AI to democratize creation for non-technical users. The discussion outlines how this approach treats AI as a tool for human expression rather than an autonomous creator, aiming to counteract the homogenization of global culture and narrow demographic biases in current tech development. Ultimately, the event advocates for a future of education focused on cultivating taste and self-knowledge, allowing artists to leverage AI to raise creative ceilings while maintaining authentic individual agency.
- a16z39 min
Text to Video: The Next Leap in AI Generation
Anjney Midha, Andreas Blattmann, Robin Rombach
Released on November 21st, Stable Video Diffusion is a state-of-the-art open-source model that generates short video clips from single input images by prioritizing the learning of complex physical properties like 3D consistency and camera movement. The architecture employs diffusion methodology over autoregressive methods to optimize perceptual details and utilizes LoRA adapters for scalable control of camera motion, while training strategies focused on temporal dynamics and specific 3D orbit refinement to achieve surprising reasoning capabilities in as few as 2,000 iterations. This release continues the team's philosophy of driving innovation through algorithmic efficiency rather than sheer compute volume, having already catalyzed a rapid ecosystem of community experimentation and set a roadmap for longer sequences and future audio integration.