acceptodds
Under review as a conference paper at ICLR 2027

TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models

Abstract

Reward models are a core bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision-based feedback, which cannot be fully satisfied by manually crafted reward functions or task-specific human annotations. Existing open-source VLM reward judges such as RoboReward adopt a simple 1–5 trajectory progress scoring scheme, lacking pairwise preferences essential for modern RLHF, DPO and Bradley-Terry frameworks, while also failing to optimize agents' video scene understanding directly. Augmenting RoboReward with pairwise comparison and video-QA supervision causes inherent inconsistency between pairwise preferences and pointwise scores, introducing training noise and hurting downstream performance—an issue that aggregation methods such as TrustJudge cannot resolve. To address this, we propose TrustRoboReward, a multi-paradigm reward modeling framework equipped with the Preference-Ordered Isotonic Score Editing (POISE) module. We construct a unified four-paradigm dataset containing trajectory progress scoring (Score-A), video-QA answer quality scoring (Score-B), and their pairwise preference counterparts (Pair-A, Pair-B). We find pairwise labels align better with human judgment than pointwise scores, inspiring us to calibrate pointwise scores to avoid score-pair reversals against pairwise preferences. The POISE module rectifies pointwise scores and eliminates cross-paradigm reversal conflicts unresolved by TrustJudge. Theoretically, POISE provably reduces training-corpus score-pair reversal conflicts from 20.15% to 0%, whereas TrustJudge still retains 20.46% reversal conflicts on the same training corpus. Evaluated on our multi-paradigm benchmark, the Qwen3-VL-4B model trained with POISE achieves an overall reward score of 77.96%, nearly matching GPT-5-mini (78.09%, gap 0.13%) and outperforming the strongest RoboReward-4B baseline by 10.13%. It also lifts test-time score-pair consistency to 71.90%, exceeding both RoboReward-4B (57.26%) and GPT-5-mini (68.09%). Further integrating TrustJudge aggregation during inference boosts our model's overall score to 78.57%, surpassing the proprietary GPT-5-mini teacher model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.