Can We Trust the Judge? Building Reliable Evaluation for Language Models
Abstract
LLM-as-a-Judge systems are now embedded in training pipelines, safety evaluation, and production deployments at scale—yet the field lacks a shared framework for assessing whether these evaluators actually measure what they intend to, in the contexts where they are deployed. JUDGe addresses this gap by reframing evaluation validity as a systems problem rather than a measurement problem: how does a judge's error profile interact with what is upstream and downstream of it, and what are the consequences when it fails in context? We bring together NLP researchers, ML systems practitioners, alignment scientists, and industry evaluators around a shared taxonomy of judge failure modes—surface sensitivity, sycophancy, criteria drift, positional bias, reasoning chain validity, safety-relevant meaning preservation, and inter-judge correlation—and their cascading interactions in live pipelines. The workshop will produce a community-facing Judge Deployment Disclosure Template, analogous to a model card but for evaluation infrastructure. Through contributed papers, junior spotlights, structured debate, and a practitioner panel, JUDGe will crystallize open problems and lay groundwork for evaluation standards the field currently lacks.