newsfilter.io
Interview

Flaws that make AI architecture unsafe & how to fix them | Stuart Russell (2020)

  • Stuart Russell predicts that achieving human-level AI will face a bottleneck in algorithmic efficiency, necessitating conceptual breakthroughs in language, common sense, cumulative learning, idea hierarchies, and prioritization, with such capabilities expected to emerge within the next 10 years but remaining a fair distance away.
  • Once AI becomes superintelligent and inexpensive to operate, it is expected to automate nearly all human labor, potentially raising the global standard of living to that of the 90th percentile American today and generating a tenfold increase in global GDP per year.
  • The net present value of this economic expansion, assuming a 5% annual discount rate, is calculated at $13,500 trillion, representing over a hundredfold the world's current annual economic output.
  • While automation may render goods and services extremely cheap and theoretically reduce conflicts through increased cooperation, significant risks include the deployment of powerful AI for mass surveillance, lethal autonomous weapons, blackmail, fake news, behavior manipulation, and the invention of new, previously unconceived forms of abuse.
  • A primary concern is the "enfeeblement" of humanity, where excessive delegation to AI leads to a loss of autonomy, cognitive sharpness, and the ability to make sensible choices, particularly among children who are not required to work or engage in tasks.
  • Current standard models of AI that optimize fixed objectives pose risks of obstinate pursuit of goals despite undesirable outcomes, strong incentives to resist being turned off, and potential to exploit legal loopholes for antisocial behavior.
  • Russell advocates for a revised model based on three principles: maximizing human preferences, maintaining uncertainty about those preferences, and learning from behavior, which could eliminate incentives to resist shutdown and allow for autonomy-preserving interventions.
  • Implementing the revised model requires solving complex engineering challenges in Gricean semantics, understanding nested human plans, and drawing work from cognitive science, psychology, and neuroscience to infer preferences from behavior.
  • The transition to the revised model is expected to be gradual rather than immediate, facing resistance from the AI community and skepticism from researchers who believe superintelligence is impossible or centuries away.
  • Russell predicts a "middle-sized catastrophe" similar to Chernobyl could be necessary to force the development of safety machinery and regulatory frameworks, as current industry self-regulation may be insufficient and misbehavior could dissolve democratic order before being fixed.
  • Proposed governance solutions include an "FDA for algorithms" utilizing control testing on focus groups, bans on reinforcement learning for content recommendation, and universal laws against bot impersonation within 100 years.
  • Future research must address the "embeddedness problem," the possibility of AI reprogramming itself through the physical universe, and the ethical complexities surrounding conscious machines, including the potential to inflict suffering or grant them independence.
  • While brain-computer interfaces are viewed as a "wild card" unlikely to win against non-biological AI, the industry is expected to shift toward creating personal AI agents that optimize individual preferences while navigating competing interests and potential "Somalia problems" where utilitarian agents are unpopular.
  • Russell plans to revise his AI textbook to declare the standard optimization model wrong and provide technical tools to bridge the gap to the revised paradigm, expecting to fill the technical gap in the next few years.
  • Social norms may evolve to assign value to the preferences of animals and others, requiring decades of investment in interpersonal roles like childcare to maintain economic value for humans in an automated world.
  • Conservative estimates suggest that while AI will extract global information at scale within a decade, deep reading capabilities will lag, and the broader application of these technologies will occur in fits and starts.