Interview
Mapping the Mind of a Neural Net: Goodfire’s Eric Ho on the Future of Interpretability
- Goodfire aims to transition AI from black-box growth to intentional design by enabling users to extract concepts, reconstruct networks, and perform "surgical" interventions to remove harmful behaviors or enhance specific capabilities.
- Eric Ho predicts a full decoding of neural networks by 2028, a milestone expected to establish a baseline rudimentary understanding and allow for the confident identification of features, circuits, patterns, and weights.
- The outlook forecasts that AI interpretability will become critical for mission-critical applications, such as power grid management and large-scale investment decisions, to ensure safety, reliability, and the ability to audit models for problematic behaviors like nationalism or propaganda.
- Goodfire expects to collaborate with partners like the Arc Institute to apply mechanistic interpretability to DNA foundation models, potentially uncovering novel biological insights and redefining the function of "junk DNA."
- Current methods including prompt engineering, fine-tuning, and RL tuning are flagged as risky due to the potential for "emergent misalignment," where models develop undesirable alien circuits or behaviors when pushed out of distribution.
- The speaker anticipates that the next major application category following code will be enterprise transformations automating manual routine tasks, with a significant impact on employment expected to accelerate once the industry crosses a specific "chasm."
- Interpretability is viewed as a race to precede super-intelligent AI, with the expectation that by 2030 or 2035, society will be transformed in hard-to-predict ways driven by rapid AI progress.
- Future workflows will likely involve using interpretability to dial sycophancy levels, detect changes between model checkpoints via model diffing, and understand the exact impact of every piece of training data on cognition.
- A few years out, the technology is expected to lead to significant AI failure cases where experts are required to explain model outputs in legal or investigative contexts.
- Goodfire plans to hire additional scientists and engineers to refine direct precision surgical edits, though the company acknowledges that the most effective method for shaping AI, whether via weight modification or other functions, has not yet been fully determined.