Fireside Chat, Interview, Conference Presentation
Vipul Ved Prakash, Together AI | RAISE Summit 2025
- Together AI, co-founded and led by Paul Ved Prakash, has operated as an "AI acceleration cloud" for nearly three years, focusing on optimizing large-scale model training and inference through hardware-software co-engineering.
- The company recently launched its first Grace Blackwell 200 cluster in Memphis, delivering 1.4 petaflops of computing power on a single rack to facilitate massive supercomputer-like training computations.
- Hardware evolution has shifted rapidly from consumer-grade GPUs (e.g., NVIDIA 3090s) to specialized data center supercomputers where all participating machines function as a single computational unit with ultra-fast interconnectivity.
- To address generative AI hallucinations, the industry is adopting "reasoning" and "test time compute" mechanisms that allow models to self-correct before outputting responses.
- Model efficiency is improving via techniques that activate only small subsets of parameters during inference, enabling real-time token processing with larger models.
- Together AI's platform hosts over 200 open-source models, allowing customers to utilize them directly, fine-tune them, or apply reinforcement learning to adapt behavior for specific use cases.
- Approximately 80% of Together AI's current business revenue is generated by "AI-native" companies, many of which report weekly growth rates of roughly 10%.
- Enterprise adoption is driven by a need for specialized infrastructure to manage high thermal loads, high power density per rack, and the requirement for orchestration and observability layers that standard cloud architectures lack.
- Strategic priorities include maximizing hardware utilization by dynamically mixing training, inference, and reinforcement learning workloads to optimize cost performance on expensive silicon.
- The market trend is shifting from "SaaS" to "Agentech," a progression requiring increased horsepower for multi-step reasoning and complex application workflows.