Interpretability as a Science: Toward Rigorous Foundations for Understanding LLMs
Abstract
LLM interpretability has attracted increasing attention, driven by questions about what concepts LLMs encode, where they are represented, and how those representations give rise to behavior. These are the right questions, but the field has yet to converge on what valid answers look like or what kind of evidence can establish them. Thus, despite rapid progress, interpretability remains epistemologically fragile. This reflects the need to scientifically ground interpretability research. Interpretability needs what other sciences already have: a principled account of hypothesis formation and testing, formal criteria for what counts as a genuine explanation, rigorous notions of measurement and falsifiability, and experimental designs that can distinguish mechanisms from artifacts. This workshop brings together researchers from diverse disciplinary background spanning interpretability, causal representation learning, neuroscience, mathematics, statistics, and physics to discuss how to establish interpretability as a rigorous empirical science. It will foster structured dialogue with researchers from adjacent fields that have confronted analogous problems to identify frameworks and experimental practices that can be adapted to interpretability. Through keynote talks and facilitated breakout discussions, we will diagnose what interpretability gets right and what it gets wrong, and work toward shared criteria for how its claims can be formalized, tested, and refined, working toward a more scientifically rigorous foundation for interpretability.