Medical Reasoning with Multimodal Foundation Models
Abstract
Vision-language foundation models have rapidly advanced pattern recognition in medicine, yet they remain fundamentally weak at reasoning, connecting observations to knowledge through explicit, verifiable, multi-step inference under uncertainty. This gap is the critical bottleneck separating today's models from trustworthy clinical decision support. Med-Reasoner is a one-day workshop at NeurIPS 2026 that convenes machine learning researchers, multimodal and vision-language model experts, medical AI scientists, and practicing clinicians around four concrete research threads: (i) faithful chain-of-thought for diagnosis and its verifiability; (ii) multimodal grounding of visual evidence in clinical and biological knowledge; (iii) evaluation frameworks measuring reasoning quality beyond answer accuracy, including faithfulness, grounding, and calibration; and (iv) probabilistic and agentic reasoning for sequential decision-making. The workshop builds on the first edition held at CVPR 2026 (30 submissions, 14 accepted, full room capacity, Amazon Health Services sponsor) and is redesigned for the NeurIPS audience around reasoning as a foundational ML problem. A structured panel debate and community open-problems synthesis will produce a shared agenda for standardized reasoning evaluation metrics, a concrete deliverable the field currently lacks. Estimated attendance is 100–150, drawing from the VLM, multimodal learning, medical imaging, computational pathology, and clinical AI communities.