Conference Presentation, Interview, Fireside Chat
Matt Hicks, Red Hat | RAISE Summit 2025
- Event Context: TheCUBE is covering RAIDS Conference 2025 in Paris, France, focusing on AI infrastructure, software, app development, and agent building across enterprise and consumer sectors.
- Market Sentiment: John Furrier notes Red Hat's phenomenal performance over the last three years, citing significant contribution to the growth of IBM stock and overall market success.
- Core Themes: Discussions at the conference center on the convergence of mainstream cloud, on-prem distributed computing, sovereign AI, and sovereign cloud, driven by the intersection of open source and AI.
- Key Product Launches: Red Hat has announced multiple recent and near-term releases, including:
- Red Hat Enterprise Linux for Developers and RHEL for business developers.
- AI Inference Server and Red Hat AI Expanded Support for multiple models with accelerated protocols.
- OpenShift Lightspeed GA and Advanced Developer Suite.
- In-Vehicle OS approaching GA.
- Strategic Partnerships: Red Hat is strengthening its leadership ecosystem through partnerships with AMD, NVIDIA, Meta, Google Cloud, Azure, and Oracle.
- Market Readiness: Enterprises are actively adopting AI but have not yet reached "full throttle" deployment; open source is identified as the primary driver to scale this transition from experimentation to production.
- Cost Strategy: Red Hat's infrastructure strategy prioritizes minimizing the cost per AI token call to prevent enterprise budgets from exploding as agentic workloads introduce multi-step reasoning and increased token usage.
- Technology Focus: Red Hat advocates for smaller, open-source models and efficiency techniques to make unit prices low enough for scalable enterprise use.
- Product Solution: The Red Hat Inference Server, built on VLLM, is designed to run any open-source model on any GPU provider to solve utilization issues.
- Problem Addressed: Enterprises often achieve only ~20% GPU utilization (e.g., on AMD or NVIDIA cards) without optimization.
- Target Outcome: The solution aims to increase GPU utilization to 90-100% to maximize hardware investments.
- Scaling Architecture: Red Hat is developing LLMD to enable scaling inference workloads across clusters, allowing models to be run on diverse architectures as a unified unit.
- Paradigm Shift: Matt Hicks draws a parallel between historical server evolution and the current AI landscape:
- Past: CPU servers running Linux with middleware and Kubernetes orchestration.
- Present/Future: GPU servers running RHEL AI or Red Hat Inference Server with LLMs as middleware.
- Innovation Drivers: Industry experts suggest hardware performance is nearing a plateau, identifying software innovation as the primary area for future step-up functions and efficiency gains.
- Power Constraints: Power consumption is identified as a bounding function for hardware; software solutions must manage workloads efficiently across x86 estates and AI models that coexist.
- Technical Mechanisms: Software innovation will focus on new forms of memory management and KV cache capabilities to unlock enterprise capabilities previously reserved for larger entities like OpenAI.
- Open Source Engagement: Key projects for developers to explore include VLLM (for technical implementation) and LLMD (for cluster management).