Fireside Chat, Interview
Together AI & SemiAnalysis: In Conversation Together AI's Vision For the Future of AI Infrastructure
- Hyperscalers face potential obsolescence due to architectural path dependence, with networking performance described as significantly inferior to new deployments even two generations ahead.
- The cloud computing market, currently valued at $1.2 trillion, is projected to triple in size over the next decade, driven largely by generative AI's central role.
- A completely new technology stack is expected to be built from the ground up to address the distinct silicon, workloads, services, pricing, and packaging requirements of generative AI.
- Imminent achievement of model capability thresholds is anticipated to unlock numerous downstream use cases, increasing demand for both inference and training workloads.
- The ClusterMax benchmarking initiative, active for roughly one year, continues evaluating dozens of providers across network performance, security, storage, reliability, and GPU health metrics.
- Transitioning to Blackwell GPU architectures may result in existing optimized code performing slower or only marginally faster (30%) unless kernels are rewritten to leverage features like asynchrony and mixed precision.
- Users renting initial Blackwell GPUs may experience no speed advantage over previous generations without specific hardware-software co-design optimization.
- Inference workloads are shifting toward massive scale deployment, with architects increasingly designing models optimized for inference time rather than just training.
- Sparse, wide mixture-of-experts architectures are driving system requirements for high batch sizes and low-latency communication across hundreds of GPUs, necessitating workload aggregation across massive clusters.
- Serving efficiency for leading open-source models is expected to reduce costs from $8 to approximately 30 cents per million tokens, enabling previously unviable use cases.
- Reinforcement learning environments are transitioning beyond human-created data, allowing models to learn business workflows in custom sandboxes independent of internet data.
- Enterprises are expected to require full-service capabilities for GPU rental, training, optimization, and serving, as most lack the internal expertise to execute these complex workflows independently.
- A completely new set of applications and workloads not yet invented or optimized is expected to emerge within the next year.