Grounded User Simulation for Model Evaluation and Training:\\ Diversity, Fidelity, and Validity
Abstract
User simulation is becoming core infrastructure for evaluating and training interactive AI systems, from conversational search and recommender systems to agentic tool-use benchmarks and post-training pipelines. Yet the field lacks shared standards for when simulated users are reliable proxies for real people, especially across languages, dialects, demographic groups, tasks, and deployment contexts. This workshop will bring together researchers in agent evaluation, information retrieval, dialogue systems, HCI/social simulation, reinforcement learning, recommender systems, synthetic data, multilingual NLP, and fairness to study user simulation through three lenses: diversity of represented users and scenarios, fidelity to individual and population-level behavior, and validity of conclusions drawn from simulation. The workshop will combine invited talks, contributed papers, posters, a debate on whether simulated users can be trusted, and a cross-community breakout. As Year-1 outputs, we will seed shared infrastructure around TREC-style conversational search and τ-family agentic tool-use settings, including a census-conditioned baseline simulator, reporting checklist, validation/failure-mode taxonomy, and benchmark catalog.