MIB: A Mechanistic Interpretability Benchmark

Aaron Mueller; Atticus Geiger; Sarah Wiegreffe; Dana Arad; Iván Arcuschin; Adam Belfki; Yik Siu Chan; Jaden Fiotto-Kaufman; Tal Haklay; Michael Hanna; Jing Huang; Rohan Gupta; Yaniv Nikankin; Hadas Orgad; Nikhil Prakash; Anja Reusch; Aruna Sankaranarayanan; Shun Shao; Alessandro Stolfo; Martin Tutek; Amir Zur; David Bau; Yonatan Belinkov

MIB: A Mechanistic Interpretability Benchmark

Authors	Aaron Mueller Atticus Geiger Sarah Wiegreffe Dana Arad Iván Arcuschin Adam Belfki Yik Siu Chan Jaden Fiotto-Kaufman Tal Haklay Michael Hanna Jing Huang Rohan Gupta Yaniv Nikankin Hadas Orgad Nikhil Prakash Anja Reusch Aruna Sankaranarayanan Shun Shao Alessandro Stolfo Martin Tutek Amir Zur David Bau Yonatan Belinkov
Publication date	2025
Journal	Proceedings of Machine Learning Research
Event	42nd International Conference on Machine Learning
Volume \| Issue number	267
Pages (from-to)	45069-45108
Number of pages	40
Organisations	Interfacultary Research - Institute for Logic, Language and Computation (ILLC)
Abstract	How can we know whether new mechanistic interpretability methods achieve real improvements? In pursuit of lasting evaluation standards, we propose MIB, a Mechanistic Interpretability Benchmark, with two tracks spanning four tasks and five models. MIB favors methods that precisely and concisely recover relevant causal pathways or causal variables in neural language models. The circuit localization track compares methods that locate the model components—and connections between them—most important for performing a task (e.g., attribution patching or information flow routes). The causal variable track compares methods that featurize a hidden vector, e.g., sparse autoencoders (SAE) or distributed alignment search (DAS), and align those features to a task-relevant causal variable. Using MIB, we find that attribution and mask optimization methods perform best on circuit localization. For causal variable localization, we find that the supervised DAS method performs best, while SAEs features are not better than neurons, i.e., non-featurized hidden vectors. These findings illustrate that MIB enables meaningful comparisons, and increases our confidence that there has been real progress in the field.
Document type	Article
Note	Proceedings of the 42nd International Conference on Machine Learning, 13-19 July 2025, Vancouver Convention Center, Vancouver, Canada.
Language	English
Published at	https://proceedings.mlr.press/v267/mueller25a.html (Final published version)
Other links	https://github.com/aaronmueller/MIB
Downloads	mueller25a (Final published version)
Permalink to this page

Back

UvA-DARE

Digital Academic Repository

MIB: A Mechanistic Interpretability Benchmark