newsfilter.io
Conference Presentation, Product Demonstration

Zero Flaps, Zero Excuses: Network Reliability at AI Scale | Credo x TensorWave | RAISE 2026

  • Executive Leadership & Company Overview

    • Phil Kuhman (SVP, Global Sales, Credo) and Jeff Tatarchuk (Co-founder, Chief Growth Officer, Tensorwave) led the discussion.
    • Credo is a publicly traded connectivity solutions company with a current valuation of ~$1.3 billion, projected to grow to $2.4 billion according to analysts.
    • Tensorwave is an AMD-exclusive "Neocloud" focused on scalable GPU infrastructure.
  • Tensorwave Scale & Strategic Positioning

    • Currently operates the largest deployment of AMD GPUs: 8,192 MI325s live in production.
    • Data center capacity spans Arizona, Pennsylvania, and Florida, with secured long-term power agreements totaling up to two gigawatts in North America.
    • Backed by AMD Ventures and Magnetar Capital (the fund that guided CoreWeave from seed to IPO).
    • Strategic focus on AMD due to its optimized memory architecture for high-throughput inference and GenAI workloads compared to NVIDIA competitors.
  • The Reliability Imperative in AI Infrastructure

    • Traditional copper active electrical cables (AECs) are limited to 7-meter distances, necessitating optical solutions for longer reaches in large-scale clusters.
    • AI training sessions across hundreds of thousands of GPUs require zero link flaps; a single link failure forces a rollback to the previous checkpoint, incurring massive time and compute costs.
    • Oracle initially faced a 95% "No Trouble Found" (NTF) rate on early H100 clusters due to fabric reliability issues with standard third-party optics.
    • Credo's engagement with Oracle reduced deployment time from six weeks to under one week, minimizing idle GPU factories and preventing billions in lost revenue (based on an estimated $4/GPU-hour).
  • Credo's Zero-Flap Transceiver Technology

    • Developed specifically to meet Oracle's requirement for optical transceivers 1,000x more reliable than standard commodity parts.
    • Achieves this reliability through a six-point differentiation strategy:
      • Hardening: Hardware undergoes rigorous thermal cycling in 20 chambers in Taiwan to break and fix links via firmware iteration.
      • In-Band Telemetry: Monitors all six links simultaneously to detect root causes (e.g., dust, fiber crimping, ESD damage) rather than just reporting errors.
      • Remote Diagnostics: Enables firmware updates and health checks without relying solely on switch-side diagnostics.
      • Health Scoring: Generates Red/Yellow/Green metrics based on Bit Error Rates (BER) and Signal-to-Noise Ratio (SNR) to predict degradation.
      • Preventative Action: Automatically flags marginal links for removal before they cause fabric disruptions.
    • Credo holds an estimated 88% market share and is driving industry standardization through the OCP Optics Reliability Workstream.
  • Architectural Strategy: "Copper Where You Can, Optics Where You Must"

    • Active Electrical Cables (AECs) replace fiber at the rack host for lower power consumption, higher reliability, and zero link flap risk (limited to 7 meters).
    • Zero-Flap optics are reserved for Top-of-Rack (ToR) spine connections where physical reach exceeds 7 meters.
    • Credo has trademarked "Purple Cables" to denote active electrical cables in deployed networks.
  • Future Outlook & Industry Collaboration

    • Credo is co-creating the Zero-Flap Optics specification with Oracle to establish new industry baselines for AI networks.
    • Proposes adding 1 PPS (Pulse Per Second) signals to the ZF Optics spec for enhanced network synchronization.
    • System vendors are currently pre-qualifying Zero-Flap optics for compatibility with all network operating systems.
    • Forward-looking statement: Reliable networking is becoming "table stakes" for AI networks; the industry must move beyond technology designed 20 years ago.
    • Strategic goal for Neoclouds: Reduce the time-to-revenue for GPU clusters by solving network reliability bottlenecks that currently cause idle infrastructure.