Interview, Conference Presentation
Vipul Ved Prakash, Together AI | RAISE Summit 2025 1
Company Milestones & Definition
- Together AI is approaching its third anniversary, having been founded approximately three years ago.
- The company operates as an "AI acceleration cloud" facilitating large-scale model training and inference workloads.
- Core technology relies on hardware-software co-optimization to maximize GPU computing efficiency.
- Together AI positions itself as a dedicated cloud provider for generative AI, distinct from traditional CPU-based data centers.
Hardware Infrastructure & Evolution
- Hardware capabilities have shifted dramatically in three years, moving from gaming-grade GPUs (e.g., Nvidia 3090) to massive supercomputer clusters.
- The company recently launched its first Grace Blackwell 200 cluster in Memphis.
- This specific cluster delivers 1.4 petaflops of computing power within a single rack.
- Modern data centers function as unified supercomputers where all machines participate in single training computations.
- Hardware design progress has accelerated interconnectivity between servers and chips, enabling near-instant data transfer speeds.
Software Capabilities & Open Source Strategy
- The platform currently hosts 200 open-source models ready for immediate deployment.
- Users can utilize models directly, fine-tune them, or apply reinforcement learning for specific use cases.
- Inference capabilities now include "reasoning and test-time compute," allowing models to self-correct and reduce hallucinations.
- Model efficiency has improved, with inference processes activating only a small subset of parameters to handle higher token throughput in real-time.
Market Dynamics & Customer Demographics
- Approximately 80% of Together AI's business is derived from "AI-native" companies.
- Key customer sectors include generative media, robotics, and new language model designers.
- Downstream enterprise applications utilizing generative AI include customer support, healthcare, back-office automation, and data processing.
- Customer growth rates for AI-native clients are reported at approximately 10% weekly.
- Enterprise demand is shifting from blank-check spending to focusing on price-performance metrics and end-to-end IP ownership.
Technical Challenges & Operational Requirements
- AI cloud architectures differ fundamentally from traditional data centers regarding power density per rack and silicon types.
- High thermal loads in AI hardware necessitate a robust "default tolerance layer" to manage failure rates.
- Enterprise adoption requires orchestration and observability tools to manage workloads effectively.
- To maximize utilization of expensive hardware, systems must dynamically mix and match workloads (e.g., running training or RL during low inference load periods).
Future Outlook & Industry Trends
- The availability of supercomputing-as-a-service is expected to lower barriers and stimulate an increase in startups.
- The market is transitioning from the "year of the killer app" (inference) to a focus on multi-step reasoning and agent technologies.
- Generative AI is viewed as the impending "digital mechanism of the world," driving massive CapEx investment and enterprise adoption.
- There is a significant gap in the enterprise market for integrated stacks that combine hardware with necessary software orchestration layers.