acceptodds
Under review as a conference paper at ICLR 2027

PolyPRM: Mandate-Conditioned Process Evaluation for Role-Specific Agents

Abstract

Process reward models (PRMs) typically evaluate intermediate reasoning and actions against a single objective. For role-specific agents, however, the same candidate process may have different values under different decision responsibilities. We introduce RIVA-Bench, a matched benchmark that holds decision contexts and candidate slates fixed while varying task-defined mandates and selector-independent, provenance-tracked utility targets. RIVA-Bench contains 111,805 labeled candidate processes, 12,872 decision states, and 50 roles across eight tasks. We also introduce PolyPRM, a 1.5B mandate-conditioned evaluator that learns across roles while preserving their distinct candidate orderings. PolyPRM achieves the highest direct-selection score on all eight tasks and outperforms frontier LLM judges supplied with the same complete mandate. It improves frozen ReAct and Debate agents over their base selection policies in all 16 evaluated settings and attains the best or tied-best non-oracle outcome in 15. Counterfactual ablations show that its selections follow mandate semantics rather than role names. Joint training further matches or exceeds routed role-specific experts across all six datasets while using one checkpoint and \(1/r\) of their aggregate training data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.