Conference Presentation
Keynote by Sunghyun Park, CEO & Co-Founder of Rebellions | RAISE Summit 2026
- Sunghyun Park, CEO and co-founder of Seoul-based chip startup Rebellion, cited a conviction in Korea's semiconductor talent, SK Hynix memory ecosystem, and Samsung Foundry as the catalyst for establishing an AI chip company in Asia rather than the US.
- Rebellion's strategic investors and validating partners include SK Hynix, ARM, Samsung Foundry, SK Telecom, and Korea Telecom.
- The company has moved beyond prototyping to mass production, delivering over 100 server racks to customers and operating multi-megawatt NPU clusters within weeks of deployment.
- Hardware optimization targets a single metric: "tokens per second per watt per dollar," prioritizing inference efficiency over absolute peak performance in a commoditized market.
- The custom silicon utilizes a hybrid SRAM and HBM memory-centric architecture featuring:
500 MB of on-chip SRAM powered by a unique mesh network delivering >190 TB/s bandwidth to support decode stages.
- Four HBM3e units totaling 144 GB capacity to handle larger models, long input sequences, and extensive knowledge bases during pre-fill.
- The chip design incorporates a scalable chiplet architecture allowing for non-wafer-tape additions of CPU or I/O die; a recently taped-out I/O die achieves up to 1.0 TB/s chip-to-chip bandwidth.
- The processor operates at approximately 600 TDP, enabling air-cooled data center deployment without requiring liquid cooling infrastructure upgrades.
- Rebellion delivers a full-stack AI infrastructure platform, including custom silicon, servers, racks, and scale-up/scale-out solutions.
- Comparative production data from SK Telecom indicates the platform achieves:
- 3x lower Capital Expenditure (CapEx) per rack compared to same-class NVIDIA platforms.
- 2.8x higher tokens generated per dollar.
- 2.2x more tokens processed per watt of power.
- Software strategy prioritizes native support for open-source ecosystems (PyTorch, Hugging Face, VLLM, Red Hat) over proprietary stacks, accepting <5% performance compromise to ensure broad compatibility and commoditization.
- The software stack leverages PyTorch 2.0 with
torch.compileand VLLM as abstraction layers to decouple end-user applications from underlying hardware specifics. - The system is currently running 24/7 for mission-critical applications at SK Telecom (ADOT), handling up to 50 million phone calls daily with variable traffic loads up to batch sizes of 32 and 64.
- Deployment locations include SK Telecom in South Korea, ARM collaborations in Saudi Arabia (Humane, Aramco), and the first financial sector reference in Korea with KB Bank.
- Rebellion is collaborating with ARM to develop a hybrid CPU/NPU orchestration layer to support complex Generative AI (AGI) workloads within a single chip or system architecture.
- The platform currently processes an average of 14 million API calls daily, supporting MaaS (Model as a Service) via Kubernetes-integrated deployments.
- Park emphasizes the need for supply chain diversification and cost optimization, positioning Rebellion as a heterogeneous compute alternative to NVIDIA's dominance in both training and inference domains.