newsfilter.io
Fireside Chat, Conference Presentation, Panel

Operating AI Infrastructure at Multi-Model, Multi-Cloud Scale | RAISE Summit 2026

  • Strategic Shift to Inference-First Models:

    • Nebius is pivoting from traditional IaaS (managed Kubernetes, VMs) to a managed "Token Factory" service, recognizing that while training is an investment, inference is the primary driver of business growth and AI application.
    • Customers increasingly prefer abstraction layers where they specify traffic patterns (RPS, tokens) rather than raw infrastructure metrics (GPU counts, megawatts, data center specs).
    • Base10 describes its market positioning as "inference is everything," addressing the gap between physical GPU infrastructure and the need for reliable, economic, scaled model serving.
  • Infrastructure & Deployment Strategies:

    • Base10 (Aggregation Model):
      • Operates across 18–20 cloud providers spanning 90+ regions to aggregate demand and secure GPU supply, including H100 clusters.
      • Provides multi-cloud orchestration to satisfy data residency, latency, and redundancy requirements, effectively creating a global compute mesh.
      • Solves "multi-model performance" and "cohesive developer experience" as key value pillars beyond raw hardware access.
    • Nebius (Vertical Integration Model):
      • Leverages deep integration into its own data centers and hardware design (e.g., custom server RAM, PCB debugging, specific cooler optimization) to optimize cost-to-performance ratios.
      • Utilizes in-house expertise to debug issues at the physical layer (hardware, network) rather than just application code, enabling faster incident resolution.
      • Currently serves as a provider for Base10, illustrating a potential for industry collaboration despite competitive positioning.
  • Market Trends & Model Dynamics:

    • Open Source Acceleration:
      • Adoption of open-source models is driven by their ability to cover frontier-level capabilities in coding and planning, reducing the need for proprietary training on standard tasks.
      • Companies utilize a "hybrid chef" strategy: using open-source models as a base, fine-tuning them with proprietary data to outperform frontier models in niche domains, or integrating specialized external models (e.g., Pionote for diarization).
      • A maturity curve exists where customers start with closed-source APIs (OpenAI, Anthropic) but transition to fine-tuned open-source models as costs scale and proprietary data advantages become critical.
    • Token Efficiency:
      • Benchmark performance is increasingly tied to "reasoning budget" (tokens processed); a cheaper per-token model (e.g., GLM 5.2) may be less efficient per task than a more expensive frontier model (e.g., Opus) if it requires more tokens to solve.
      • Specialization allows for novel experiences where models answer questions faster and with higher token efficiency than off-the-shelf alternatives.
  • Commercial Models & Pricing:

    • Nebius Token Factory:
      • Pricing is flexible and customer-centric; while the underlying infrastructure uses GPU hours, the service communicates value in "tokens" to match customer traffic expectations.
      • For large enterprise clients, pricing remains transparently based on GPU hours to allow clear economic modeling, though the platform supports token-based abstraction.
      • The value proposition focuses on transparency, allowing customers to trade cost for redundancy or varied hardware configurations.
    • Base10:
      • Primarily charges per token where natural, but retains flexibility to translate usage into other units (e.g., GPU hours) based on customer needs.
      • Heavy investment in auto-scaling infrastructure ensures customers pay only for active usage rather than reserved capacity, mitigating "token maxing" waste.
  • Geopolitics & Sovereignty:

    • The market is projected to remain fragmented regarding model sovereignty, with ongoing shifts between US, Chinese, and European (e.g., French open-source) models.
    • Open-source models are expected to persist as a permanent fixture to ensure availability and mitigate geopolitical risks associated with closed-source providers.
    • Base10 and Nebius anticipate continued demand for localized or sovereign model hosting as data residency laws and political controversies (e.g., US government restrictions on Anthropic) evolve.