Conference Presentation, Panel, Fireside Chat
WEKA, Cohere, Nebius: Rewriting the Rules Agentic AI and the Race to Scale
RAISE SummitLauren Vaccarello, Cécile Robert-Michon, Dan Shtan, VYACHESLAV TYKHONOVSKYI, DAN GALPIN, Danila
- Future AI infrastructure demands a complete re-architecting to support large-scale models, driven by the need for extreme speed, agility, and flexibility to match a high-velocity development cycle where the landscape may shift within two weeks or five years.
- Agentic AI is projected to trigger a massive spike in inference demand and usage due to longer sessions and complex interactions, requiring infrastructure to optimize both quality and efficiency as tokens become smarter during scaling.
- Hardware and software strategies prioritize maximum flexibility, with virtualized platforms targeting 99.7% of NVIDIA benchmark standards and occasionally reaching 104%, allowing customers to avoid locking into hardware that may be obsolete within two weeks.
- Platform architecture aims for complete cloud agnosticism across all training stack layers, supporting deployments on every major cloud provider as well as private and air-gapped environments to ensure data control and sovereignty.
- Enterprise adoption faces significant hurdles in regulated sectors like finance, healthcare, manufacturing, public sector, and energy, where security, data privacy, and reticence toward new technology create barriers to rapid implementation.
- Current cost models prioritize immediate system capability and functionality over financial efficiency, with optimizations and cost reductions anticipated only in later stages of adoption as software evolves faster than hardware.
- Storage and performance benchmarks show a converged solution using NVMe disks or spare CPU cores can deliver up to 20 times greater checkpoint performance improvements compared to traditional object storage, leveraging existing network infrastructure.
- Scalability plans must accommodate a wide spectrum of users ranging from individuals with limited GPU budgets to clusters with thousands of interconnecting GPUs, while legacy software stacks present unresolved questions regarding the integration of legacy AI systems.
- Organizational planning is characterized by a clear focus on immediate, short-term goals (approximately two weeks) and a recognized lack of visibility into quarterly or longer-term requirements, necessitating the abandonment of rigid, long-term fixed plans.
- Resource constraints will limit the abundance of available capacity, forcing reliance on maximizing current assets and utilizing creative software engineering to unlock performance gains that hardware alone cannot achieve.