Executive Summary
* The Google DeepMind mechanistic interpretability team has made a strategic pivot over the past year, from ambitious reverse-engineering to a focus on pragmatic interpretability:
* Trying to directly solve problems on the critical path to AGI going well [[1]]
* Carefully choosing problems according to our comparative advantage
* Measuring progress with empirical feedback on proxy tasks
* We believe that, on the margin, more researchers who share our goals should take a pragmatic approach to interpretability, both in industry and academia, and we call on people to join us
* Our proposed scope is broad and includes much non-mech interp work, but we see this as the natural approach for mech interp researchers to have impact
* Specifically, we’ve found that the skills, tools and tastes of mech interp researchers transfer well to important and neglected problems outside “classic” mech interp
* See our companion piece for more on which research areas and theories of change we think are promising
* Why pivot now? We think that times have changed.
* Models are far more capable, bringing new questions within empirical reach
* We have been disappointed by the amount of progress made by ambitious mech interp work, from both us and others [[2]]
* Most existing interpretability techniques struggle on today’s important behaviours, e.g. they involve large models, complex environments, agentic behaviour and long chains of thought
* Problem: I...
Make this part of your paper trail
Save this paper to a shelf, write a review, and keep your own notes.