acceptodds
Under review as a conference paper at ICLR 2027

Task-Directed Reward Shaping with Weight-Conditioned Policies

Abstract

Reward shaping has been used successfully in large-scale reinforcement learning systems across a wide-range of domains. Yet poorly chosen reward weights can direct learning toward shaping signals at the expense of the task objective. Selecting these weights often requires manual iteration, repeated policy training, or bilevel optimization. We introduce Mix-Max Policy Search (MMPS), which takes a multi-objective reinforcement learning perspective to tune a fixed library of auxiliary rewards for a given task. MMPS trains a shared, weight-conditioned actor–critic across reward mixtures and uses its predicted primary-task value to select the mixture for each rollout. The selected mixture guides both on-policy data collection and the PPO update, sharing learning across weights while directing training toward behaviors predicted to yield high task return. Across nine IsaacGym tasks, MMPS achieves higher task performance than existing weight-tuning methods under matched environment-interaction budgets, with its largest gains on difficult manipulation tasks, achieving the score of policies trained with expert-tuned reward weights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.