Language2Utility: Transferable Natural-Language Preference Grounding for Multi-Objective Reinforcement Learning
Abstract
Multi-objective reinforcement learning (MORL) addresses sequential decision making with competing objectives, but preference-conditioned MORL typically takes numerical preference specifications as given. In practice, however, users express desired trade-offs in natural language, creating a semantic gap between linguistic preferences and the numerical preference specifications underlying multi-objective utility. Motivated by a relational view of preference semantics over objective sets, we study transferable grounding of natural-language preferences across heterogeneous objective spaces and introduce Language2Utility. Language2Utility encodes preference-objective relations and contextualizes them with a permutation-equivariant set model over variable-cardinality objective sets. To accommodate practical preference expressions that may not uniquely specify a numerical trade-off, learning combines ordinal constraints with supervision over sets of semantically compatible numerical preferences. We further decouple shared semantic grounding from downstream domain-specific behavioral realization. A semantic-to-behavior bound formalizes the separation by relating language-conditioned utility regret to semantic misalignment and realization error. Across four MO-Gymnasium benchmarks, Language2Utility improves both semantic grounding and behavioral relevance in-domain. Under leave-one-domain-out transfer, where both the target task and its objective cardinality are unseen during training, Language2Utility achieves better zero-shot transfer than the compared baselines. Human-authored evaluations across three practical application scenarios further support semantic transfer to new task contexts and objective sets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.