acceptodds
Under review as a conference paper at ICLR 2027

PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing

Abstract

Robotic policies are commonly evaluated by binary success rates, which compress an entire execution into a terminal outcome and provide limited information about the execution process. We present PRM-as-a-Judge, a process-level evaluation paradigm that uses Process Reward Models (PRMs) to estimate dense task progress from robot rollout observations and represent each execution as a progress curve. Based on this curve, we introduce the OPD (Outcome–Process–Diagnosis) metric system, which organizes trajectory evaluation around three sequential questions: (1) how far the execution reaches (Outcome), (2) how efficiently it advances (Process), and (3) how progress deteriorates through regression or stagnation (Diagnosis). To evaluate progress judges before policy auditing, we introduce RoboPulse, an interval-level benchmark that measures their ability to identify whether task progress increases or decreases over temporal intervals. Experiments across diverse tasks and data sources show a clear advantage of specialized PRM judges over general-purpose vision-language model (VLM) judges in task-progress assessment. Finally, we apply PRM-as-a-Judge to large-scale policy auditing in both simulation and real-world rollouts, revealing: (1) stage-wise bottlenecks, (2) differences in successful execution quality, and (3) distinct regression and stagnation patterns that are not captured by terminal success rates. We will release RoboPulse and the complete PRM-as-a-Judge evaluation toolkit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.