Interview
Who's trying to steal AI models? And what could they do with them? | Sella Nevo
- AI models are projected to have significant national security implications in the very near future, specifically regarding their potential to assist in developing bioweapons capable of killing millions, hacking critical infrastructure, or being utilized by terrorist organizations, rogue states, and anarchistic hacker groups.
- In the near future, AI's influence on biology is described as incredibly plausible, tangible, and close, with specialized models already driving the state of the art in biological research.
- Predictions for the future indicate that neural networks may be inverted to reveal model weights, a process similar to how hash functions were eventually broken, as the field of distillation advances and becomes less opaque.
- Future trends suggest the market size for AI may outpace model size growth, potentially lowering barriers for model extraction, while quantized models may require less information to remain useful to attackers.
- Model extraction attacks are expected to become more common and feasible, potentially allowing attackers to infer 10 trillion bits of information, with distributed attacks across public APIs capable of extracting weights within a few months.
- Internal network access is anticipated to enable attackers to query models faster and with less monitoring than public APIs, while task-specific distillation may allow the creation of usable models for narrow tasks with less data.
- The "80,000 Hours" organization plans to expand its advising scope, hire a head of video and a head of marketing, deploy a yearly marketing budget of three million dollars, and maintain a transcript collection available on its site.
- The flood forecasting system developed at Google Research is expected to reduce flood injuries and costs in Africa and Asia by directly notifying individuals via Android notifications, collaborating with humanitarian organizations like the Red Cross, and alerting governmental authorities for evacuations.
- Confidential computing is forecasted to become a widely deployed industry standard with minor overheads, though it requires trusted execution environments, separate chips, and enhanced hardware support from companies like NVIDIA to function effectively as a security measure.
- Future AI security strategies will likely involve "defense in depth" to exponentially decrease attacker success rates, alongside the adoption of hardened interfaces such as pre-approved code and output rate limitations for isolated networks.
- Red teaming is predicted to evolve to simulate higher-resource attackers by granting diverse privileges and skill sets in cyber, physical, and human intelligence domains, with external third-party teams necessary to provide reliable security signals.
- The demand for AI security professionals with expertise in cyber and physical security is expected to grow significantly, becoming a key bottleneck for organizations as threats expand to include hundreds of different attack vectors.
- Human intelligence attacks in the future are expected to involve grooming employees by eroding boundaries or leveraging ideological beliefs, as well as extortion tactics threatening family members in specific geopolitical regions.
- Zero-day vulnerabilities are predicted to be found daily in hundreds of products, with nation-states potentially cheating by acquiring these vulnerabilities through mandatory reporting channels to be used by offensive cyber organizations.
- Supply chain attacks remain a serious problem requiring security measures across all linked organizations, while side channel attacks involving temperature, electricity usage, noise, and electromagnetic radiation (Tempest) will continue to pose significant risks in cloud environments.
- Air-gapped networks will remain vulnerable to USB-based malware unless specific physical and procedural controls are implemented, and the gap between AI capabilities and security measures could widen without comprehensive security benchmarks.
- The growth of the AI security sector may see better outcomes where increased resource investment significantly reduces the number of organizations capable of attacking or lowers the likelihood of success.
- Specific operational constraints, such as the "80,000 Hours" application taking 10 minutes and the need for audio engineering teams to master and technical edit episodes, are planned for future execution.
- While confidential computing is a critical step forward, it is not viewed as a silver bullet and will require additional measures to address all remaining attack vectors.