newsfilter.io
Fireside Chat, Interview

Together AI & SemiAnalysis: In Conversation Together AI's Vision For the Future of AI Infrastructure

  • Hyperscalers face potential obsolescence due to architectural path dependence, with networking performance described as significantly inferior to new deployments even two generations ahead.
  • The cloud computing market, currently valued at $1.2 trillion, is projected to triple in size over the next decade, driven largely by generative AI's central role.
  • A completely new technology stack is expected to be built from the ground up to address the distinct silicon, workloads, services, pricing, and packaging requirements of generative AI.
  • Imminent achievement of model capability thresholds is anticipated to unlock numerous downstream use cases, increasing demand for both inference and training workloads.
  • The ClusterMax benchmarking initiative, active for roughly one year, continues evaluating dozens of providers across network performance, security, storage, reliability, and GPU health metrics.
  • Transitioning to Blackwell GPU architectures may result in existing optimized code performing slower or only marginally faster (30%) unless kernels are rewritten to leverage features like asynchrony and mixed precision.
  • Users renting initial Blackwell GPUs may experience no speed advantage over previous generations without specific hardware-software co-design optimization.
  • Inference workloads are shifting toward massive scale deployment, with architects increasingly designing models optimized for inference time rather than just training.
  • Sparse, wide mixture-of-experts architectures are driving system requirements for high batch sizes and low-latency communication across hundreds of GPUs, necessitating workload aggregation across massive clusters.
  • Serving efficiency for leading open-source models is expected to reduce costs from $8 to approximately 30 cents per million tokens, enabling previously unviable use cases.
  • Reinforcement learning environments are transitioning beyond human-created data, allowing models to learn business workflows in custom sandboxes independent of internet data.
  • Enterprises are expected to require full-service capabilities for GPU rental, training, optimization, and serving, as most lack the internal expertise to execute these complex workflows independently.
  • A completely new set of applications and workloads not yet invented or optimized is expected to emerge within the next year.