Convention Grounding: Language as a Coordination Prior for Open Coordination Points
Abstract
Connected autonomous vehicles from different manufacturers must coordinate at open coordination points, a four-way intersection at simultaneous arrival, or a merge where each stack assumes a different rule, where local information fails to resolve a single order; rules frozen in each proprietary stack leave humans no way to alter which convention governs a mixed fleet. We introduce convention grounding: an operator broadcasts a convention as a natural-language sentence, and each receiving vehicle combines that text with its own observation to produce its role-differentiated action, under a reward that never observes the convention. In SUMO, a continuous traffic simulator, we provide, to our knowledge, the first controlled causal demonstration of this mechanism: freezing the scene and flipping only the convention text flips the agent's committed action (flip rate 1.00, zero collisions), while swapping the projected vectors of the two conventions maintains the flip rate but inverts correctness. At a four-way intersection, four independently trained policies, sharing no gradients, policy, training episodes, or projection parameters, coordinate through a common convention signal and their own observations. Generalization to the tested held-out phrasings fails: the role distinction remains unrecovered from the projected representation, and runtime selection is shown only within the grounded convention set. No large generative model is invoked during driving: encoding each broadcast costs one forward pass through a frozen encoder.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.