Conference Presentation, Keynote
Ramine Roane, CVP @ AMD: Unlocking the Next Wave of AI Open Foundations for Intelligent Systems
- AI intelligence is projected to scale logarithmically with total compute across pre-training, post-training, and inference phases as the industry transcends limits imposed by publicly available data, shifting toward reasoning models that generate multiple chains of thought via extremely verbose outputs.
- Post-training for verifiable domains like math and coding will standardize reinforcement learning through compile-and-verify workflows, while future inference serving will likely utilize pre-fill and decode disaggregation where initial KV caches are stored on shared disks to enable low-latency generation across many GPUs.
- AMD anticipates its MI400 GPU will launch in 2026 alongside the MI350, which will support full-scale racks already being deployed by certain CSPs, while the MI355x, described as AMD's Blackwell equivalent, launched concurrently with NVIDIA's Blackwell architecture.
- AMD HPM memory is expected to deliver a 2.8x performance advantage over the Grace Blackwell NVL72 in HBM capacity, with claims of superior performance relative to Blackwell on FP8 workloads due to NVIDIA's TensorRT-LLM lacking native support at the time of benchmarking.
- AMD hardware is forecast to support AI inference across diverse sectors including laptops, embedded systems, automotive, satellites, Mars rovers, and medical equipment, with data center CPUs and GPUs expected to serve major entities like OpenAI, Meta, and X.AI for ChatGPT, Llama, and Grok inference.
- An enterprise AI software stack utilizing Kubernetes, VLLM, SGLang, and LLMD with telemetry is planned for release in the second half of this year for enterprise and sovereign nations, leveraging Silo AI's acquisition to facilitate international collaboration with European labs and nations.
- A massive data center project with a capacity approaching one gigawatt is under construction in Grenoble, France, based on AMD GPUs, while the broader market trend favors sovereign nations adopting compute infrastructure from at least two vendors to mitigate vendor lock-in regarding pricing and delivery schedules.
- Open-source ecosystems, OCP racks, Ultra Ethernet, and UL links are identified as critical standards for the future of AI, with a roadmap indicating that non-human intelligence will eventually compete with and surpass human intelligence, mirroring the global adoption scale of the Industrial Revolution.
- Developers scanning provided code are eligible for several hours of free GPU compute with potential for renewal, and the provider asserts full compatibility with Hugging Face and PyTorch workflows currently possible on other GPUs to ensure optimal utilization in large-scale inference environments serving tens of thousands to hundreds of thousands of customers.