newsfilter.io
Conference Presentation

Keynote by Sunghyun Park, CEO & Co-Founder of Rebellions | RAISE Summit 2026

  • Sunghyun Park, CEO and co-founder of Seoul-based chip startup Rebellion, cited a conviction in Korea's semiconductor talent, SK Hynix memory ecosystem, and Samsung Foundry as the catalyst for establishing an AI chip company in Asia rather than the US.
  • Rebellion's strategic investors and validating partners include SK Hynix, ARM, Samsung Foundry, SK Telecom, and Korea Telecom.
  • The company has moved beyond prototyping to mass production, delivering over 100 server racks to customers and operating multi-megawatt NPU clusters within weeks of deployment.
  • Hardware optimization targets a single metric: "tokens per second per watt per dollar," prioritizing inference efficiency over absolute peak performance in a commoditized market.
  • The custom silicon utilizes a hybrid SRAM and HBM memory-centric architecture featuring:
    • 500 MB of on-chip SRAM powered by a unique mesh network delivering >190 TB/s bandwidth to support decode stages.

    • Four HBM3e units totaling 144 GB capacity to handle larger models, long input sequences, and extensive knowledge bases during pre-fill.
  • The chip design incorporates a scalable chiplet architecture allowing for non-wafer-tape additions of CPU or I/O die; a recently taped-out I/O die achieves up to 1.0 TB/s chip-to-chip bandwidth.
  • The processor operates at approximately 600 TDP, enabling air-cooled data center deployment without requiring liquid cooling infrastructure upgrades.
  • Rebellion delivers a full-stack AI infrastructure platform, including custom silicon, servers, racks, and scale-up/scale-out solutions.
  • Comparative production data from SK Telecom indicates the platform achieves:
    • 3x lower Capital Expenditure (CapEx) per rack compared to same-class NVIDIA platforms.
    • 2.8x higher tokens generated per dollar.
    • 2.2x more tokens processed per watt of power.
  • Software strategy prioritizes native support for open-source ecosystems (PyTorch, Hugging Face, VLLM, Red Hat) over proprietary stacks, accepting <5% performance compromise to ensure broad compatibility and commoditization.
  • The software stack leverages PyTorch 2.0 with torch.compile and VLLM as abstraction layers to decouple end-user applications from underlying hardware specifics.
  • The system is currently running 24/7 for mission-critical applications at SK Telecom (ADOT), handling up to 50 million phone calls daily with variable traffic loads up to batch sizes of 32 and 64.
  • Deployment locations include SK Telecom in South Korea, ARM collaborations in Saudi Arabia (Humane, Aramco), and the first financial sector reference in Korea with KB Bank.
  • Rebellion is collaborating with ARM to develop a hybrid CPU/NPU orchestration layer to support complex Generative AI (AGI) workloads within a single chip or system architecture.
  • The platform currently processes an average of 14 million API calls daily, supporting MaaS (Model as a Service) via Kubernetes-integrated deployments.
  • Park emphasizes the need for supply chain diversification and cost optimization, positioning Rebellion as a heterogeneous compute alternative to NVIDIA's dominance in both training and inference domains.