newsfilter.io
Conference Presentation, Product Demonstration

Rethinking AI Infrastructure for the Age of Reasoning | Jeehoon Kang, FuriosaAI | RAISE 2026

  • Strategic Context & Mission

    • AI inference workloads are driving a surge in new data center construction, with power availability identified as the primary industry bottleneck.
    • Furiosa AI's mission is to establish sustainable AI computing by delivering high-performance chips and systems that maximize token generation per watt.
  • Product Launch: Renegade (2nd Gen Chip)

    • Production Timeline: Mass production of the second-generation Renegade chip begins this year with a target output of 20,000 units.
    • Manufacturing Specs: Fabricated using TSMC 5nm process technology with approximately 40 billion transistors and a die size near the reticle limit.
    • Memory Architecture: Features a hybrid HBM-SRAM design integrating 2x HBM dies (48GB total) and 256MB of SRAM to optimize power efficiency.
    • Power Specifications: Operates at a 180W TDP, allowing a single server to consume only 3kW and enabling 5–6 servers per standard air-cooled rack.
    • Performance Gains: Delivers 30% to 100% higher token output than competitors within the same power-constrained server environment.
    • Target Market: Designed specifically for traditional data centers reliant on air cooling rather than direct liquid cooling.
  • Technical Architecture: Tensor Contraction Processor (TCP)

    • Design Philosophy: Built from first principles of tensor operations rather than modifying existing GPU architectures, aiming to resolve the trade-off between power efficiency and programmability.
    • Hardware Structure: Utilizes a multi-slice architecture where data is uniformly distributed across SRAM and processed via a tensor-native pipeline.
    • Deterministic Performance: Eliminates dynamic features (e.g., dynamic frequency scaling, complex cache coordination) to ensure predictable kernel latency, unlike GPUs.
    • Optimization Strategy: Employs a global search compiler that leverages predictable latency and a tractable search space to select optimal kernel implementations.
    • Search Space Management: Reduces optimization complexity through two high-level abstractions:
      • Shapes: Declarative definitions for memory mapping (e.g., uniform SRAM distribution).
      • Tactics: Declarative descriptions of instruction selection and scheduling that mirror the hardware's tensor unit pipeline.
  • Software & Ecosystem

    • Programming Models: Supports a unified compiler stack allowing users to mix high-level abstractions (for ease of use) and low-level languages (for precise hardware control) within a single program.
    • Developer Tools: Provides OpenAPI-compatible API endpoints, a Vision-Language Model (VLM) compatible inference engine, and Kubernetes device plugins.
    • Deployment Readiness: Products are currently available for immediate order with delivery timelines measured in days.
    • Live Demo: Showcased at Booth 22, integrating the VoxTraw and Mistral models to translate speech into eight languages in real-time.
  • Future Roadmap: Torque (3rd Gen Chip)

    • Target Market: Designed for hyperscale systems utilizing direct liquid cooling and significantly higher power envelopes.
    • Scalability: Expected to scale performance 10x to 30x compared to the Renegade generation.
    • Partnership: Co-developed in partnership with Broadcom.
  • Executive Announcements

    • Furiosa AI CEO Jun Paik will present "Powering AI Infrastructure for the 21st Century's Manhattan Project" on the Master Stage tomorrow at 1:40 PM.