Physical Understanding for Decision-Making: Bridging Foundation Models and Reliable Agents
Abstract
Classical control and planning approaches—such as model predictive control, trajectory optimization, and physics- based simulation—addressed this challenge by encoding physics and mathematical constraints explicitly, and achieved remarkable precision in structured environments. However, these methods do not scale gracefully to the diversity, partial observability, and contact-richness of real-world deployments, and they require painstaking manual modeling of every relevant physical relationship. The recent emergence of foundation models for decision-making—vision-language-action (VLA) models, robot foundation models, agentic systems, and interactive world models—offers a compelling alternative. These systems have sharply improved long-horizon video generation, embodied policy learning, and multimodal action grounding, building on a long tradition of world models for prediction and planning. Yet most of these models are still trained and evaluated primarily for visual realism or short-term prediction accuracy. For embodied agents, this gap is critical: a robot that cannot reason about the outcomes of its actions will fail on contact-rich manipulation, long-horizon planning, and safety-critical deployment. The field now stands at an inflection point where foundation and world models are powerful enough to serve as the backbone of physical agents, yet the scientific community lacks shared definitions, evaluation protocols, and benchmarks for physical understanding as a distinct and measurable capability. This workshop is designed to fill that gap before incompatible paradigms and vocabularies become entrenched.