Task-Directed Reward Shaping with Weight-Conditioned Policies
Abstract
Reward shaping has been used successfully in large-scale reinforcement learning systems across a wide-range of domains. Yet poorly chosen reward weights can direct learning toward shaping signals at the expense of the task objective. Selecting these weights often requires manual iteration, repeated policy training, or bilevel optimization. We introduce Mix-Max Policy Search (MMPS), which takes a multi-objective reinforcement learning perspective to tune a fixed library of auxiliary rewards for a given task. MMPS trains a shared, weight-conditioned actor–critic across reward mixtures and uses its predicted primary-task value to select the mixture for each rollout. The selected mixture guides both on-policy data collection and the PPO update, sharing learning across weights while directing training toward behaviors predicted to yield high task return. Across nine IsaacGym tasks, MMPS achieves higher task performance than existing weight-tuning methods under matched environment-interaction budgets, with its largest gains on difficult manipulation tasks, achieving the score of policies trained with expert-tuned reward weights.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.