Interpretability for Discovery: Understanding and Discovering Novel Knowledge in AI Models
Abstract
We propose a workshop on Interpretability for Discovery to explore how model interpretability can support the discovery and understanding of novel knowledge in AI systems. The workshop aims to build a community around three core questions: (i) How can existing interpretability methods be adapted to handle architectures, biases, and data modalities that differ from the LLM setup? (ii) How can interpretable model representations be translated into novel knowledge and discoveries? (iii) Which open questions can model interpretability help answer? Our goal is to bring together researchers from machine learning and related disciplines to foster discussion on methodologies, experimental design, and real-world applications so that we can enable ground-breaking discoveries from insights encoded in AI systems.