Workshop on the Linguistic Principles for Foundation Models
Abstract
How a task is expressed, including its wording, structure, and notation, is not merely the channel through which we query foundation models (FMs); it shapes what they can do. The same problem rendered in different but meaning-equivalent forms, whether a paraphrase, another natural language, or a formal notation such as code or logic, can shift reasoning accuracy dramatically and induce entirely different internal representations, even with model parameters held fixed. The linguistic and symbolic medium is therefore a design axis for FM capability, comparable to architecture and scale, rather than a neutral interface. This workshop establishes the principled study of this medium as a first-class research agenda, linking classical linguistic structures such as compositionality, ambiguity, tokenization, pragmatics, and typological variation to model reasoning, generalization, and alignment. It convenes machine learning researchers and linguists to advance a direction complementary to scaling: improving FMs by rethinking the media through which they perceive, reason, and act.