DROPS: Robust Pareto Set Learning under Preference Drift for Multi-Objective Robotic Control
Abstract
Multi-objective Robotic Control is increasingly important in real-world applications, with Multi-Objective Reinforcement Learning (MORL) providing a promising approach for learning policies over conflicting objectives. However, existing MORL methods typically assume clean observations and fixed environment dynamics. These assumptions often fail in practice, where uncertainty can cause severe **_preference drift_**, substantially undermining the robustness of learned control policies. In this paper, we propose **DROPS**, a Preference-Drift-Aware approach for learning a Robust Pareto Set of control policies against uncertainty. Specifically, DROPS constructs **_(i) a drift-aware trust region_** for each preference, introduces **_(ii) hierarchical adversarial perturbations_** to establish a degradation bound with a policy-independent coefficient, and leverages **_(iii) contrastive drift alignment_** to recover robust preference-to-policy mappings. Central to our method is strengthening the robustness of the learned policy Pareto Set by tightening the degradation bound and aligning drifted preferences, enabling DROPS to learn a Pareto Set robust to preference drift. Experiments on five continuous robotic-control tasks, together with a discrete benchmark, show its consistently superior performance in obtaining robust control policies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.