Conference Presentation, Interview, Fireside Chat
Matt Hicks, Red Hat | RAISE Summit 2025
- Token usage is projected to increase by an order of magnitude due to agentic workloads and multi-step reasoning, creating a capability ceiling for enterprises if costs are not reduced.
- Strategy involves prioritizing smaller open source models and mathematical efficiency to minimize costs per question while scaling across clusters via the LLMD project to handle increased inquiry volumes.
- Hardware utilization goals aim to shift from typical 20% GPU usage to 90% or 100% by deploying the Red Hat Inference Server on any GPU provider with VLLM.
- The industry architecture is expected to transition from CPU-centric middleware like Kubernetes to unified distributed GPU systems where purpose-tuned software operates as integral units.
- Deep-sea software innovation is identified as the primary driver of future advancement, as hardware performance approaches the limits of its step-up function.
- Power constraints are anticipated to act as a bounding function, necessitating infrastructure estates where x86 systems manage workloads that interface with models on alternative hardware.
- Software advancements in memory management, KV cache, and cluster learning will enable enterprises to achieve capabilities comparable to larger entities like OpenAI.
- Massive infrastructure demand will be driven by the emergence of agents and the integration of reinforcement learning with human feedback into multi-step reasoning processes.