Interview, Fireside Chat
How To Build Generative AI Models Like OpenAI's Sora
- Future AI models are projected to simulate real-world physics, fluid dynamics, and biological systems with accuracy potentially surpassing billion-dollar NOAA weather models, enabling applications in drug discovery, gene therapy, and engineering calculations that replace obsolete Fortran-based CAD kernels.
- Generative capabilities are expected to expand into lip-syncing accuracy for languages like Hindi, direct EEG-based stroke prediction and thought reading, and the creation of high-fidelity AI replicas of individuals using as little as one hour of video footage without massive training datasets.
- Sora is anticipated to eliminate the need for physical cars by simulating carless societies and complex vehicle dynamics, though current iterations may exhibit specific physics errors such as floating objects, incorrect lane driving, or disjointed structures.
- The development of "World Models" and space-time architectures is expected to shift from black-box deep learning to explainable foundation models, with synthetic data from game engines like Unreal Engine becoming a standard resource for self-training models and reducing reliance on massive external data collection.
- Founders from non-AI backgrounds are expected to reach the cutting edge of the field within six to nine months of intensive study, enabling the creation of foundation models during YC Winter 24 batches with as little as $500,000 and dedicated GPU access.
- Startups are predicted to iterate a hundred times faster than previous rates through dedicated GPU clusters, allowing companies to compete with industry giants in verticals like hardware design, software copilots, and protein modeling without requiring hundreds of millions in funding.
- Specific technological advancements include quadratic reductions in runtime complexity for EEG data processing, low-resolution video compression techniques reducing data requirements, and the use of high-quality textbook data to allow smaller models like GPT-2.5 to outperform larger models in constrained tasks.
- Risks and misconceptions include the "straight line path" to success, the historical dead-end of reinforcement learning for early robotics, and the theoretical "mosquito drinking its own blood" problem regarding synthetic data, though recent developments suggest synthetic data can effectively generate its own training loops.
- Industry evolution is expected to move past the necessity of billion-dollar data centers for foundational models, with companies like Playground and Sonato competing with Midjourney and Midjourney using fewer resources, while others like Kscale Labs and Draft8 focus on deploying consumer humanoid robots and replacing traditional CAD software.
- Historical context indicates that the Visual Transformer (2020) and World Model (2018) concepts remain foundational, while current trajectories suggest a shift where 10 trillion parameter models like Sora will drive long-term visual consistency and temporal accuracy in video generation.