Interview, Fireside Chat
How Real-Time AI Video Is Changing How Creators Work
- A large consumer AI moment is anticipated soon, driven by a model reaching quality and cost thresholds sufficient for novel social experiences, with the immediate next one to two months dedicated to enhancing controllability for professional studio use cases.
- Output reliability for lip synchronization, motion controls, and camera directions is targeted to increase from a current 80–90 percent range to 99.9 percent.
- Hollywood studios are expected to increase AI usage by 10x to 100x in the coming months once legal and data residency obstacles are resolved, potentially adopting workflows combining LLMs in Blender with H3 Max for close to 100 percent controllability.
- The market is predicted to follow a pattern of brief quiet periods followed by simultaneous bursts in base models, latency, and controllability, with the Generative Media Conference next week expected to shift focus toward Hollywood and new AI studios.
- New model versions include "H3 Max Turbo," offering five-second video generation in 1.5 seconds at half the cost of standard versions with minor quality trade-offs, and "H3 Max," designed as the default model for other platforms due to superior speed and efficiency.
- The "H3 Max" architecture is projected to become the industry's only model capable of generating continuous, action-controlled video up to 60 minutes long, featuring a two-minute memory span and an evolving system prompt for coherence.
- Efficiency improvements will be pursued through specialized hardware inference pipelines, post-training capabilities for both open and closed-source models, and optimization of prompt expansion, diffusion, and VAE components.
- Industry-wide optimizations and kernel engineering aim to achieve 70–80 percent theoretical model efficiency (MFU), while hardware upgrades from Hoppers to Blackwell GPUs are expected to reduce wall clock time by 2x to 3x.
- Plans include deploying a "Fall Live" website for continuous video demonstration with user voting, integrating with social channels like Instagram and TikTok using AI IP holders for style LoRAs, and exploring voice prompting tools for real-time director interaction.
- The ecosystem will expand with additional LoRA fine-tunes for specific styles, camera angles, and lip syncing, while new infrastructure will move away from one-off training runs toward service-based offerings.
- Future consumer applications involve running video models on home hardware and creating real-time interactive experiences where video displays update based on conversation, though optimizations may not translate perfectly.
- Compute constraints are expected to persist, driving the need for token efficiency to support "token market fit" where individual creators can spend thousands of dollars or 10k+ tokens monthly on video generation.
- The "H3 Max" model is positioned as the first next-generation open-source video model capable of accepting references, served in parallel across eight GPUs to maintain efficiency without significant communication overhead.