Conference Presentation, Product Demonstration
Rethinking AI Infrastructure for the Age of Reasoning | Jeehoon Kang, FuriosaAI | RAISE 2026
Strategic Context & Mission
- AI inference workloads are driving a surge in new data center construction, with power availability identified as the primary industry bottleneck.
- Furiosa AI's mission is to establish sustainable AI computing by delivering high-performance chips and systems that maximize token generation per watt.
Product Launch: Renegade (2nd Gen Chip)
- Production Timeline: Mass production of the second-generation Renegade chip begins this year with a target output of 20,000 units.
- Manufacturing Specs: Fabricated using TSMC 5nm process technology with approximately 40 billion transistors and a die size near the reticle limit.
- Memory Architecture: Features a hybrid HBM-SRAM design integrating 2x HBM dies (48GB total) and 256MB of SRAM to optimize power efficiency.
- Power Specifications: Operates at a 180W TDP, allowing a single server to consume only 3kW and enabling 5–6 servers per standard air-cooled rack.
- Performance Gains: Delivers 30% to 100% higher token output than competitors within the same power-constrained server environment.
- Target Market: Designed specifically for traditional data centers reliant on air cooling rather than direct liquid cooling.
Technical Architecture: Tensor Contraction Processor (TCP)
- Design Philosophy: Built from first principles of tensor operations rather than modifying existing GPU architectures, aiming to resolve the trade-off between power efficiency and programmability.
- Hardware Structure: Utilizes a multi-slice architecture where data is uniformly distributed across SRAM and processed via a tensor-native pipeline.
- Deterministic Performance: Eliminates dynamic features (e.g., dynamic frequency scaling, complex cache coordination) to ensure predictable kernel latency, unlike GPUs.
- Optimization Strategy: Employs a global search compiler that leverages predictable latency and a tractable search space to select optimal kernel implementations.
- Search Space Management: Reduces optimization complexity through two high-level abstractions:
- Shapes: Declarative definitions for memory mapping (e.g., uniform SRAM distribution).
- Tactics: Declarative descriptions of instruction selection and scheduling that mirror the hardware's tensor unit pipeline.
Software & Ecosystem
- Programming Models: Supports a unified compiler stack allowing users to mix high-level abstractions (for ease of use) and low-level languages (for precise hardware control) within a single program.
- Developer Tools: Provides OpenAPI-compatible API endpoints, a Vision-Language Model (VLM) compatible inference engine, and Kubernetes device plugins.
- Deployment Readiness: Products are currently available for immediate order with delivery timelines measured in days.
- Live Demo: Showcased at Booth 22, integrating the VoxTraw and Mistral models to translate speech into eight languages in real-time.
Future Roadmap: Torque (3rd Gen Chip)
- Target Market: Designed for hyperscale systems utilizing direct liquid cooling and significantly higher power envelopes.
- Scalability: Expected to scale performance 10x to 30x compared to the Renegade generation.
- Partnership: Co-developed in partnership with Broadcom.
Executive Announcements
- Furiosa AI CEO Jun Paik will present "Powering AI Infrastructure for the 21st Century's Manhattan Project" on the Master Stage tomorrow at 1:40 PM.