acceptodds
Under review as a conference paper at ICLR 2027

What Should Critic Learn? Rethinking Critic Design for Preference-Conditioned Multi-Objective Reinforcement Learning

Abstract

Multi-objective reinforcement learning (MORL) increasingly relies on preference-conditioned policies to adapt decisions across competing objectives, yet a basic architectural question remains unresolved: how should critics represent and estimate objective-wise values? Existing methods variously use scalarized critics, shared multi-head critics, or fully separated critics, often without a principled account of what design is appropriate. We study this question by separating two sources of critic error that are typically conflated: the loss of objective-specific information induced by scalarized supervision and the cross-objective interference induced by shared trainable representations. We first show that scalarized supervision can mask substantial objective-wise errors, even when the scalarized prediction is accurate. We then characterize the interference that remains in multi-head critics through their shared representation, deriving a local gradient-interaction result and a finite-trajectory decomposition of the prediction residual introduced by parameter sharing. Guided by our proposed theory, we first conduct a controlled toy experiment designed to directly validate the predicted behavior. We then empirically evaluate scalarized, shared multi-head, and fully split critic architectures across five representative MORL algorithms—Momba, CAPQL, GPI-LS, MORL/D, and MO-MPO—on eight continuous-control environments spanning two, three, and five objectives. Across the evaluated settings, our results show that critic architecture materially affects MORL performance, with fully split critics providing a robust default across diverse evaluated settings. More broadly, our findings establish critic architecture as a fundamental design choice in MORL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.