Conference Presentation, Fireside Chat
Dr. Walden C. Rhines, Cornami: Scalable Secure AI Deploying LLMs with Fully Homomorphic Encryption
- Quantum computers with 10,000 or more sustainable qubits are expected to necessitate a shift to post-quantum encryption, with fully homomorphic encryption (FHE) identified as a quantum-proof solution that prevents data exposure in plaintext.
- The market for encrypted information sharing is projected to evolve into a major business model where data owners rent encrypted databases and charge for decryption of query results, while companies using FHE anticipate relief from GDPR liability as data remains encrypted within data centers.
- Encrypted query mechanisms are planned to secure large language model owners and users by keeping proprietary fine-tuned data and query results encrypted during processing.
- Current FHE implementation on traditional hardware faces high computational complexity, increasing processing time from one second for plaintext to approximately 11 days.
- Kornami requires a new hardware architecture moving beyond von Neumann designs, following seven years of software optimization to enable linear scalability up to 64 million cores and deliver 16 million times faster performance than NVIDIA's single instruction multiple data approach for specific data flows.
- The Kornami architecture is projected to fit up to 930 cores in the silicon area of a single NVIDIA core, achieving a 235x advantage in tokens per transistor and consuming 20 times less power (5%) for equivalent computation levels.
- While future NVIDIA cores are expected to increase power consumption as they grow, Kornami cores are projected to consume less power as feature sizes shrink.
- Existing software using standard interfaces like PyTorch is expected to run in accelerated form on Kornami hardware without significant rewriting.
- Kornami operations are projected to reduce costs by approximately three orders of magnitude per FHE operation compared to H100 chips and offer a 10 to 20x advantage in tokens per second even without full security requirements.
- At four nanometer technology, the architecture is expected to deliver 900x performance with less than a fourth of the power, while a hypothetical replacement of Elon Musk's Colossus AI data center could reduce power consumption from 150 megawatts to 130 kilowatts.
- Gains of three orders of magnitude in power efficiency are expected to mitigate fears of global data center power capacity constraints, with real-time FHE enabling AI applications to run without modification at 20x tokens per second using 60% of the power on 16-nanometer technology.
- Beta units are scheduled to reach customers during the current quarter, with volume shipping and order taking targeted for the fourth quarter or potentially September.