AIDaR: AI Data Readiness for Scientific Discovery
Abstract
Scientific AI is moving from curated prediction tasks toward systems that work directly with experimental data, scientific knowledge, and analysis workflows. Foundation models, retrieval systems, and agents are increasingly expected to query evidence, run analyses, interpret outputs, and support follow-up decisions. For these systems to be reliable, data must remain connected to the samples, assays, protocols, measurements, workflow outputs, and contextual assumptions needed to interpret results beyond the lab or notebook where they were produced. The workshop centers on one question: which measurable properties of scientific data ecosystems predict downstream AI behavior? We use AI Data Readiness (AIDaR) to describe this measurable relationship between data ecosystems and model or agent performance. AIDaR focuses on two linked problems. First, how should scientific data be organized so AI systems can use it: from raw measurements and analysis-ready files to relational or graph-structured knowledge, multimodal biomedical measurements, workflow records, and large datasets used for model training and experimental feedback? Second, how should evaluations determine where frontier models are reliable: across biological reasoning, practical data analysis, retrieval over structured scientific knowledge, assay-specific workflows, and deployed scientific AI settings? The workshop will bring together researchers, industry scientists, and engineers building scientific data systems, foundation models, agents, biomedical evaluations, and industrial assay platforms. The program includes invited talks, contributed papers, demos, panels, and breakout groups focused on data infrastructure, practical evaluations, and lessons from scientific AI deployments.