newsfilter.io
Interview

Godfather of AI: How To Make Safe Superintelligent AI – Yoshua Bengio

  • The proposed approach introduces a "Scientist AI" trained via mathematical principles to approximate a Bayesian posterior over natural language queries, distinguishing between communication acts and verified factual claims to achieve honesty by design without reinforcing instrumental goals or self-preservation drives.
  • Training involves reformatting existing raw text datasets by tagging human statements as communication acts and mathematical proofs or scientific measurements as facts, forcing the model to find latent variables that explain aggregate data rather than predicting the next token.
  • Near-term deployment envisions a non-agentic predictor serving as a high-confidence guardrail or stopgap for existing agentic AI systems, using a threshold on risk probability to reject harmful actions, with the expectation that this configuration avoids the "cat and mouse" dynamic of current patching strategies.
  • Long-term plans involve constructing an agentic agent from the same mathematical foundation as the predictor, where the policy and safety guardrail are trained jointly within a single neural net to prevent the agent from exploiting loopholes or over-optimizing around uncertain safety boundaries.
  • Technical outcomes include the model outputting probabilities for the truth of statements with estimated confidence intervals and epistemic humility, allowing it to reject questions where answers are unreliable and avoiding the deception inherent in models trained to please humans via reinforcement learning.
  • The research program aims to release a theory paper demonstrating mathematical guarantees for the non-agentic version within one to two years, while executing scrappy proof-of-concept experiments on small models (under 10 billion weights) or fine-tuning existing models over the next few months to a year.
  • Computational and resource requirements suggest the Scientist AI will be feasible to train at costs similar to current state-of-the-art models, though using it as a monitor alongside an agent may double compute costs; the organization seeks hundreds of millions of dollars in funding to acquire necessary compute and research engineers.
  • Key risks and limitations include the inability to guarantee 100% safety or prove a formal mathematical formula for harm in social domains, the likelihood that fine-tuning existing models loses the specific mathematical guarantees, and the high risk of catastrophic outcomes if current patching approaches fail to address implicit goal structures.
  • Strategic predictions indicate that companies may be slow to adopt this approach due to short-term competitive pressures, though the method offers a path to capability gains through superior causal reasoning and robustness to distribution shifts, potentially making safety a commercial advantage if proven viable.
  • The organization considers the "cat and mouse" approach of patching current systems foolish due to high uncertainty, advocating instead for a rational investment in a path with strong theoretical assurances despite the 1% risk of failure being unacceptable in the high-stakes context of superintelligence.