Interview, Conference Presentation
Vipul Ved Prakash, Together AI | RAISE Summit 2025 1
- Together AI marks its third year anniversary to introduce an AI acceleration cloud capable of handling large-scale model training and inference.
- Generative AI is projected to evolve into the world's primary digital mechanism, necessitating immense computational horsepower and optimized GPU computations.
- Hardware design will advance to ensure rapid server and chip interconnectivity, while future models will operate more efficiently at inference by activating only a subset of parameters for increased throughput.
- Inference workloads will increasingly rely on reasoning and test-time compute to enable models to self-correct, with enterprises needing to build default tolerance layers for high-thermal-load hardware failures.
- The entrepreneurial landscape is expected to shift as developers access supercomputing as a service, likely triggering a surge in startup formation due to accessible cloud infrastructure.
- Industry operations will prioritize maximizing utilization of expensive hardware by mixing workloads, such as running training or reinforcement learning during periods when inference is inactive, even as enterprises address growing gaps in orchestration and observability.