Interview, Fireside Chat
The Quest for Community-Trained Open Source AI Models
- The organization aims to accelerate AI innovation by making technology and underlying code accessible, expecting this strategy to multiply open-source efforts.
- Research strategies will prioritize pushing boundaries with minimal compute through diverse alternatives rather than a centralized path.
- Founding members intend to create neutrally aligned models, such as the Hermes series, that allow users to direct AI personas instead of enforcing a specific helpful assistant constraint.
- The industry is expected to fully accept synthetic data, resolving previous uncertainties regarding student models eclipsing master models.
- If major models like Llama 4 remain closed due to legal or business reasons, the open-source community is anticipated to replicate these capabilities independently.
- Distro is predicted to reduce bandwidth requirements by approximately 857x in worst-case scenarios and between 2,000x to 3,000x under optimistic conditions.
- Distro-scaled model performance is expected to improve with increasing model size, with the performance differential versus the Mosaic optimizer widening over time.
- The organization plans to release the Distro paper and source code in October, followed by a submission to the ICLR conference.
- Future development phases include a full-stack tooling system to enable practical community training and experiments with asynchronous communication to optimize latency in decentralized settings.
- Immediate practical applications include enabling centralized actors with multiple data centers to utilize standard interconnects like 100 Gigabit Ethernet instead of specialized high-speed options.
- Long-term predictions suggest specialized hardware like ASICs could enable training via forward passes, potentially shifting the landscape away from general-purpose GPUs, while future smartphone usage might allow models to update via inference training.
- A 7B model trained on 4 trillion tokens is considered immediately possible with the current code, potentially requiring around 1,000 rented H100 GPUs.
- Training a model equivalent to Llama 3 (405B) via the community is estimated to take until the end of next year or later due to current engineering challenges in sharding.
- The training code will be made hardware-agnostic to support mixed hardware environments, including both Apple Silicon and NVIDIA GPUs.
- The organization expects the shift to distributed training to erode centralized training structures over years as Distro scales, potentially pressuring chip manufacturers to redesign silicon for higher VRAM capacity.
- The open-source movement is expected to gain momentum through a "SETI at Home" style aspiration, encouraging individuals to contribute compute power to a shared global AI project.