Interpretability Research Already Has a Framework for Actionability
Introduction
This post was either an anonymous submission of an interesting paper or was written by a student; full credit remains with the author (linked).
Prompted by a Chris Olah talk years ago comparing interpretability to the discovery of the microscope, this post examines the recent push for interpretability methods to be "actionable" or "pragmatic" rather than curiosity-driven — including a DeepMind proposal to evaluate methods like SAEs and steering vectors by whether they solve problems on the critical path to real objectives.