newsfilter.io
Conference Presentation, Panel, Fireside Chat

Time to First Token: The Metric That Will Define Our Era | Nscale, VAST Data & NVIDIA | RAISE 2026

  • Nscale intends to vertically integrate the AI stack by building a vertically integrated hyperscaler owning and operating all five layers: energy, chips, infrastructure, models, and applications.
  • Rod Evans aims to provide global cloud services by collaborating with partners like VAST to construct the bottom three layers of the AI stack.
  • Vast Data plans to power between four to five million GPUs worldwide, largely from NVIDIA, by developing a data services layer supporting storage, databases, and compute runtimes for RAG and agentic pipelines.
  • Service providers selling tokens as a service are expected to achieve revenue double that of hourly GPU sales, creating a path to higher margins through inference services.
  • NVIDIA is predicted to drive down the cost of token delivery by establishing "token factories" where providers sell only tokens, while maintaining a forward-looking build strategy two to three years ahead of demand.
  • Planned improvements across the stack include the Vera Rubin generation for inference, a new CPU layer, and the LPX platform dedicated inference chip to address the shift from training horsepower to inference memory bandwidth.
  • The next generation of systems is planned to be cable-less between GPUs to improve efficiency and reduce infrastructure requirements.
  • Storage systems are being integrated natively into core computer networks via technologies like CMX, fundamentally altering the memory hierarchy to prioritize inference workloads.
  • Enterprises are advised to augment existing applications with new functionality rather than rebuilding from scratch, though Vast Data intends to focus exclusively on new stacks and data while discarding legacy IT systems.
  • Executive leaders are urged to personally understand the technology to avoid delegating architecture and roadmap decisions, as relying on external roadmaps or treating AI as speculative research may lead to failure.
  • Organizations failing to build within professional clouds like Nscale or via best-practice private infrastructure face high risks of failure due to the tight timeline and difficulty of the technology.
  • For large-scale deployments such as a 20-megawatt, 10,000-GPU system costing nearly $2 billion, success depends on rapid return on investment driven by tight integration of all layers.
  • Storage is identified as a critical multiplier and a fundamental part of the AI stack, with data infrastructure cited as the primary bottleneck preventing GPU systems from reaching production.
  • Immediate organizational priorities are focused on demonstrating AI's impact on the bottom line and company returns, rather than viewing agentic AI solely as a long-term research initiative.