acceptodds
Under review as a conference paper at ICLR 2027

Pareto-Consistent Preference Learning for Traffic Signal Control

Abstract

Urban traffic signal control is inherently multi-objective, yet traffic quality is not directly available as a well-defined reward vector. Instead, it must be inferred from competing criteria such as efficiency, safety, fairness, and stability, and the preferred ordering of traffic states can reverse as preferences change. Methods based on preference-independent scalar rewards or learned scores reduce these competing criteria to a single implicit compromise. We formulate traffic signal control as multi-objective preference learning under the risk of representation collapse. Our method learns a structured vector-valued traffic-quality representation and a unified preference-conditioned policy family, rather than a preference-independent scalar score and a fixed controller. To supervise this representation, we construct deterministic, schema-constrained traffic abstractions and combine rule-based objective-specific comparisons with preference-conditioned comparisons generated offline by an LLM. We further encourage dominance consistency to preserve Pareto-relevant relational structure. Based on the learned representation, we train a single controller that can realize diverse preference-conditioned trade-offs without retraining for each preference or querying the LLM during deployment. Results on CityFlow benchmarks across multiple urban networks, unseen preferences, and shifted traffic conditions indicate competitive traffic performance and improvements in ordering preservation, preference alignment, and empirical trade-off quality within the evaluated settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.