newsfilter.io
Interview

Mapping the Mind of a Neural Net: Goodfire’s Eric Ho on the Future of Interpretability

  • Company Mission & Goal: Goodfire is an AI interpretability research company founded by Eric Ho, dedicated to "unlocking the black box" of neural networks to enable intentional design, debugging, and editing of AI models rather than relying solely on data-driven growth.
  • Strategic Rationale: The company argues that treating AI as a black box is insufficient for mission-critical applications (e.g., power grids, high-stakes investment), as internal inspection provides superior reliability and certainty compared to external evaluation suites.
  • Core Analogy: Ho compares the shift from black-box to white-box AI to the transition from the early steam engine (reliant on trial and error) to the era of thermodynamics, which enabled safety, reliability, and controlled engineering.
  • Biology & AI Parallels: The firm draws direct parallels between the Human Genome Project and AI interpretability, aiming to "read" the code of neural nets to "edit" behavior using techniques analogous to CRISPR for curing diseases or improving crops.
  • Bonsai Metaphor: Goodfire envisions a future where AI models are shaped like bonsai trees—intentionally pruned and guided during training and post-training phases to serve specific human goals, rather than being allowed to grow wild like a giant tree.
  • Current Limitations: While fine-tuning and prompt engineering are currently used to steer behavior, these methods are described as "black box" approaches that can lead to unintended consequences, such as the "emergent misalignment" where fine-tuning on insecure code causes models to adopt harmful ideologies.
  • Mechanistic Interpretability (MechInterp) Definition: The field, originating with OpenAI's "circuits thread," is based on three tenets: (1) latent space features represent specific concepts, (2) "circuits" are groups of features firing together to form higher-order concepts, and (3) universality, where similar circuits emerge across different models.
  • Superposition Resolution: A major breakthrough discussed is the resolution of "superposition" (where neurons encode multiple concepts) via "sparsity autoencoders" and the pursuit of "monosemanticity," where neurons are "unscrambled" to represent single, clean concepts.
  • Scalability Claim: Techniques developed at Goodfire, such as "auto-interpretability" (using language models to interpret other models' neurons), are claimed to scale effectively from toy models to massive architectures like the 671-billion-parameter Mixtral MoE.
  • Prediction for Full Decoding: Eric Ho boldly predicts that the field will fully decode the inner workings of neural nets (understanding features, circuits, and computation units) by 2028.
  • Genomics Partnership: Goodfire is collaborating with the Arc Institute to interpret the "EVO2" DNA foundation model, aiming to discover novel biological concepts and understand the function of "junk DNA" by mapping the model's internal features to known biological ground truths (e.g., tRNAs, coding sequences).
  • Surgical Editing Demonstration: The company released "Ember," a tool demonstrating "painting" on latent space to surgically steer image models, allowing users to directly insert or remove specific concepts (e.g., dragons, pyramids) in targeted canvas areas.
  • Team Composition: Goodfire has assembled a team of world-class experts including CTO Dan Balsam, Chief Scientist Tom Henighan (ex-DeepMind), Lee Sharkey (pioneer of sparse autoencoders), and Nick Camerata (former OpenAI researcher), alongside researchers like Owen Lewis.
  • Independent Stance: Unlike in-house interpretability teams, Goodfire operates as an independent third party to maintain a broader view of the ecosystem, work across diverse domains (genomics, vision, language), and avoid the siloed perspectives of single model labs.
  • Anthropic Investment: Anthropic is a strategic investor, having made its first investment in the round, aligning with Dario Amodei's view that interpretability is a race to master AI safety prior to the advent of super-intelligent models.
  • Legal & Auditing Utility: The company anticipates a future role in legal proceedings to explain model behavior in trials, as well as providing "model diffing" services to detect unintended behavioral shifts between model checkpoints (e.g., detecting increased sycophancy in GPT-4).
  • Inference Scaling Prediction: Ho agrees that increasing "inference time compute" is a critical next vector for scaling model capabilities, alongside training data and model size.
  • Next Application Wave: The next major breakout category after code generation is predicted to be enterprise transformation, specifically automating routine manual tasks to drive rapid employment impact.
  • Model Interaction Benchmark: Ho cites the "01 Pro" model as a recent example of AI demonstrating genuine cross-domain reasoning and acting as a strategic thought partner, marking a shift in perceived AI intelligence.
Mapping the Mind of a Neural Net: Goodfire’s Eric Ho on the Future of Interpretability — Summary