I Can’t Believe It’s Not Better (ICBINB): Failure Modes of AI in Biology
Abstract
Artificial intelligence (AI) is rapidly transforming biology, enabling progress in gene regulation modelling, cellular response prediction, and protein structure determination. Despite strong performance on curated benchmarks, AI models often fail when deployed beyond controlled experimental settings. This gap arises from fundamental properties of biological systems, including inter-individual heterogeneity, substantial measurement noise, limited labelled data, weak or confounded ground truth, and intrinsic biological complexity, such that benchmark success does not reliably translate to real-world biological reliability. The workshop I Can’t Believe It’s Not Better (ICBINB): Failure Modes of AI in Biology will bring together the machine learning and life sciences communities to systematically document, analyze, and learn from negative results and real-world failure modes across genomics, transcriptomics, structural biology, and clinical prediction. In addition, we will emphasize evaluation beyond curated benchmarks toward deployment-relevant assessment, together with methodological advances that address these limitations, including robustness under distribution shift, interpretable and mechanistic modelling, causal learning, uncertainty quantification, and adaptive learning strategies that improve reliable generalization and trustworthy deployment. These challenges directly reflect core questions in modern machine learning concerning robustness, causality, interpretability, uncertainty, and evaluation in real-world settings. By centring real-world failure analysis in biologically consequential settings, the workshop aims to clarify the gap between benchmark performance and scientific reliability, establish principled evaluation perspectives for trustworthy discovery, and advance more reliable and scientifically meaningful AI systems in biology.