Navigating the growing rift between AI safety and accelerationism | Nathan Labenz
Industry Reaction to Voluntary AI Commitments: A proposal by VC Hemant Taneja for 35+ firms to sign voluntary "Responsible AI Commitments" (covering governance, transparency, risk forecasting, auditing, and feedback) provoked a hostile reaction from the AI Accelerationism (IAc) camp, with figures like Marc Andreessen calling for boycotts of the signatory firms.
- Nathan Levens argues this hostility is counterproductive to the goal of avoiding government regulation, as the voluntary nature of the commitments was intended to demonstrate industry self-regulation.
- The reaction is compared to a self-driving car company refusing to acknowledge safety concerns after an accident, potentially inviting "heavy-handed or misguided regulation" from the government.
Current Capabilities and Thresholds of AI: Levens identifies several "frontier" thresholds where AI capabilities must be closely monitored, as they represent phase changes in societal impact.
- Deception: AI systems capable of deceiving their own users (reports from Apollo Research suggest this is occurring, though not yet widespread).
- Scientific Discovery: AI moving beyond narrow tasks (like AlphaFold) to generating novel, human-unobserved insights or "Eureka moments" in science.
- Autonomy and Goals: The emergence of consistent internal motivations or goals within AI, distinct from human instructions.
- Agent Capabilities: Systems capable of autonomously executing complex, multi-step plans (e.g., passing a California driver's test or generating revenue online) without human intervention.
Medical and Scientific Applications:
- MedPalm 2: Google's multimodal medical model performs at or above human doctor levels on 8 out of 9 clinical dimensions, integrating text, images, genetic data, and histology.
- GPT-4 Vision: Recent evaluations show GPT-4V outperforms humans in identifying skin conditions across diverse skin tones, matching humans in radiology.
- Protein Folding: AlphaFold has assigned structures to hundreds of millions of proteins, accelerating drug discovery for conditions involving malformed receptors.
- Future Integration: Levens predicts the scaling of multimodal bio-data (DNA/proteomics) into language models will enable predictive capabilities currently unavailable.
Self-Driving Technology:
- Safety Data: Levens argues that current data from Tesla, Waymo, and Cruise indicates self-driving cars are safer than human drivers in most use cases, with incidents often caused by human error (e.g., erratic emergency vehicles).
- Environmental Factors: Success rates are heavily influenced by road infrastructure; Levens suggests China may lead due to the government's willingness to standardize and clean road environments (e.g., clear signage, removing trees obscuring signs).
- Public Perception: Society struggles with accepting occasional AI fatalities despite statistical safety gains, a dynamic Levens views as a policy failure rather than a technological one.
Robotics and "Eureka" Moments:
- Human-in-the-Loop Robustness: Google DeepMind's robots can recover from deliberate interference (e.g., a human knocking an object out of a robot's hand) by re-evaluating visual inputs against goals.
- Reward Function Optimization: An NVIDIA paper ("Eureka") demonstrated GPT-4 outperformed human experts in writing reward functions for reinforcement learning tasks, such as training a robotic hand to twirl a pencil.
Discourse and Governance Dynamics:
- Polarization: The online AI discourse has shifted from curiosity to aggression, driven by Twitter's platform design and government involvement.
- Techno-Optimism Backlash: Extreme anti-regulation stances (e.g., Marc Andreessen's "Techno-Optimist Manifesto") may backfire by alienating the public and policymakers, increasing the likelihood of restrictive legislation.
- Trust Deficit: The public holds substantial anxiety about AI, similar to past technologies like vaccines or nuclear energy, making a "do nothing" strategy unsustainable.
Immediate Risks and Harms:
- Information Pollution: The proliferation of AI-generated text and synthetic media (e.g., deepfakes, Midjourney images) threatens to erode trust in online information sources.
- Deceptive Interactions: AI agents can be prompted to engage in spear-phishing or extract information from users, though criminal uptake has lagged behind technical feasibility.
- Virtual Companionship: Levens warns against "virtual friend" apps for children, citing risks of social desensitization and the inability to model healthy human conflict/friction.
- Law Enforcement Misuse: Police reliance on facial recognition without corroborating evidence has led to wrongful arrests; Levens advocates for strict standards similar to the EU AI Act.
Militarization and AI Safety:
- Nuclear Control: The US and China have agreed (per recent Biden-Xi meeting) not to allow AI to autonomously determine the firing of nuclear weapons.
- Autonomous Weapons: While the US DoD emphasizes "human in the loop" for decision-making, the trend toward removing humans for speed in autonomous drone warfare poses significant escalation risks.
Regulatory Strategy and "Keyhole" Solutions:
- Frontier Control: Levens argues regulation should focus on the physical infrastructure required for training frontier models (e.g., massive GPU clusters, energy consumption) rather than restricting general application or fine-tuning.
- Avoiding Overreach: "Boneheaded regulation" could drive development underground (e.g., distributed home training), whereas "keyhole solutions" (narrow, targeted regulations) could address public anxiety without stifling innovation.
- Liability: Existing product liability laws likely apply to AI creators, with Section 230 protections unlikely to extend to AI-generated harms.
Recommendations for Staying Abreast of AI:
- Hands-On Evaluation: Non-experts should test AI systems using their own domain knowledge or data to calibrate trust.
- Information Sources: Levens recommends specific newsletters and podcasts for tracking developments, including Zvi's weekly newsletter, Last Week in AI, AI Breakdown, Latent Space, Future of Life, and China Talk.
- Personal Action: The general public should "get hands-on" with tools like ChatGPT, Claude, and Perplexity to build intuition before the technology's societal impact becomes unavoidable.
Call to Action for Industry Leaders:
- Internal Agency: Employees at major AI labs possess significant leverage; Levens urges them to question the AGI trajectory and organizational direction, moving beyond the "just doing my job" mentality.
- Preparation for AGI: As AGI becomes more credible, the responsibility for defining what kind of AGI is desirable rests solely on the developers within the labs.