acceptodds
Under review as a conference paper at ICLR 2027

DCRL: DIRECT CONTRASTIVE REWARD LEARNING FROM CARDINAL TRAJECTORY SCORES

Abstract

Trajectory-level scores evaluate complete behaviors but not individual decisions. We study learning a standalone dense Markov reward from fixed, possibly suboptimal trajectories with cardinal scores, without transition rewards, pairwise preference labels, or additional environment interaction. We present Direct Contrastive Reward Learning (DCRL), which matches similar states across trajectories to derive two local reward constraints. Contrastive ordering compares different actions using trajectory-score differences; smoothness assigns similar rewards to similar state-action pairs. DCRL combines these constraints with trajectory-score fitting, then freezes the learned reward for any downstream reinforcement-learning algorithm that accepts scalar rewards. The ordering is assumed only as a statistical tendency across state-matched comparisons, allowing individual violations. Using 10 suboptimal demonstrations per task on several Gymnasium MuJoCo benchmarks, DCRL outperforms T-REX, R4, and Frozen RRD (a fixed-data version of RRD) in final mean SAC return on every task. It achieves 3.6–10.1 times the return of the best supplied demonstration and reaches 80.0–90.5% of true-reward SAC performance on three tasks. Ablations show that contrastive ordering improves on trajectory-score fitting alone and that smoothness provides a further gain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.