ATTRIB: Workshop on Data Attribution and Provenance
Abstract
As generative AI systems become widely deployed, a central challenge is attribution: determining how model outputs relate to data. Existing approaches span contributive attribution, which estimates the causal influence of training data on model behavior, and corroborative attribution, which identifies sources that semantically support or correspond to model outputs. While both paradigms have advanced significantly, their adoption in real-world contexts (such as copyright disputes, regulatory audits, safety investigations, and output verification) remains limited. This gap reflects a mismatch between current technical outputs, such as influence scores or semantic matches, and the forms of attribution required in practice, which must be interpretable, auditable, and grounded in identifiable sources. The rapid growth of synthetic and model-generated data further complicates attribution by blurring distinctions between training data and outputs. This workshop brings together researchers and practitioners across machine learning, law, journalism, and related domains to examine the limitations of existing attribution methods, clarify practical requirements, and identify research directions that can enable reliable, actionable attribution in deployed AI systems.