Conference Presentation, Lecture
What a GPT-7 Intelligence Explosion Looks Like | Carl Shulman
- The outlook posits that as AI capabilities advance from weaker systems, it may become feasible to extract assistance for strengthening adversarial examples and neural lie detectors.
- Scenarios are anticipated where initial systems lack malicious motivations, allowing for safe incremental development, while hostile early systems can be mitigated if their motivations are detected, experimented upon, and corrected.
- A second "saving throw" is considered plausible, enabling humans to leverage misaligned AIs to solve alignment problems faster than the AIs can overthrow humanity or hack servers, though rapid uncovering of such motivations is expected to create highly volatile situations.
- Humans may fail in alignment scenarios, but this is viewed as a second chance to evaluate outputs and maintain hard power to prevent server rooting.
- AI interactions are expected to generate a rich empirical feedback loop, allowing humans to identify AI exploits used to bypass interpretability methods even if the specific mechanics of those exploits remain partially understood.
- AI contributions are predicted to reach a threshold equivalent to additional researchers, boosting effective productivity by 50 to 100 percent.
- Effective compute doubling time, currently approximately eight months, is expected to be impacted by AI automation of software innovations such as inventing transformers, discovering Chinchilla scaling, and creating Flash Attention.
- Workforce transformation is anticipated where tens of millions of GPUs could perform the work of 40 or more existing workers, scaling the effective workforce from tens of thousands to hundreds of millions.
- This scaling is projected to immediately drive diverse discoveries and the development of tremendous technologies.
- Human-level AI is described as already deep within an intelligence explosion that must initiate with systems weaker than human-level capabilities.
- Deployment strategies may involve thousands of AI instances using voting algorithms to equal the performance of a single human worker.
- AIs are expected to utilize neural networks for deeper search, similar to AlphaGo, to offset model inefficiencies by consuming more compute.
- AI systems are predicted to design synthetic training data by identifying useful skills and generating complex curricula, a task currently deemed impractical for humans due to the sheer volume of steps required.
- As sophistication increases, AIs will better identify useful skills, eventually generating training data of higher quality than human data.
- Future self-play and tasks like unit test production will allow AIs to generate their own training signals, specifically for solving programming problems difficult for current AI models.
- At the speaker's company, the generation of billions of programming challenges is assigned to AI rather than human employees.