newsfilter.io
Conference Presentation, Interview, Fireside Chat

Matt Hicks, Red Hat | RAISE Summit 2025

  • Event Context: TheCUBE is covering RAIDS Conference 2025 in Paris, France, focusing on AI infrastructure, software, app development, and agent building across enterprise and consumer sectors.
  • Market Sentiment: John Furrier notes Red Hat's phenomenal performance over the last three years, citing significant contribution to the growth of IBM stock and overall market success.
  • Core Themes: Discussions at the conference center on the convergence of mainstream cloud, on-prem distributed computing, sovereign AI, and sovereign cloud, driven by the intersection of open source and AI.
  • Key Product Launches: Red Hat has announced multiple recent and near-term releases, including:
    • Red Hat Enterprise Linux for Developers and RHEL for business developers.
    • AI Inference Server and Red Hat AI Expanded Support for multiple models with accelerated protocols.
    • OpenShift Lightspeed GA and Advanced Developer Suite.
    • In-Vehicle OS approaching GA.
  • Strategic Partnerships: Red Hat is strengthening its leadership ecosystem through partnerships with AMD, NVIDIA, Meta, Google Cloud, Azure, and Oracle.
  • Market Readiness: Enterprises are actively adopting AI but have not yet reached "full throttle" deployment; open source is identified as the primary driver to scale this transition from experimentation to production.
  • Cost Strategy: Red Hat's infrastructure strategy prioritizes minimizing the cost per AI token call to prevent enterprise budgets from exploding as agentic workloads introduce multi-step reasoning and increased token usage.
  • Technology Focus: Red Hat advocates for smaller, open-source models and efficiency techniques to make unit prices low enough for scalable enterprise use.
  • Product Solution: The Red Hat Inference Server, built on VLLM, is designed to run any open-source model on any GPU provider to solve utilization issues.
    • Problem Addressed: Enterprises often achieve only ~20% GPU utilization (e.g., on AMD or NVIDIA cards) without optimization.
    • Target Outcome: The solution aims to increase GPU utilization to 90-100% to maximize hardware investments.
  • Scaling Architecture: Red Hat is developing LLMD to enable scaling inference workloads across clusters, allowing models to be run on diverse architectures as a unified unit.
  • Paradigm Shift: Matt Hicks draws a parallel between historical server evolution and the current AI landscape:
    • Past: CPU servers running Linux with middleware and Kubernetes orchestration.
    • Present/Future: GPU servers running RHEL AI or Red Hat Inference Server with LLMs as middleware.
  • Innovation Drivers: Industry experts suggest hardware performance is nearing a plateau, identifying software innovation as the primary area for future step-up functions and efficiency gains.
  • Power Constraints: Power consumption is identified as a bounding function for hardware; software solutions must manage workloads efficiently across x86 estates and AI models that coexist.
  • Technical Mechanisms: Software innovation will focus on new forms of memory management and KV cache capabilities to unlock enterprise capabilities previously reserved for larger entities like OpenAI.
  • Open Source Engagement: Key projects for developers to explore include VLLM (for technical implementation) and LLMD (for cluster management).