acceptodds
Under review as a conference paper at ICLR 2027

RubricGuard: Rubric-Driven Falsifiable Auditing and Behavioral Intervention for Generative Reward Models on Long-Horizon Tasks

Abstract

Generative reward models and the broader LLM-as-a-judge paradigm have become the de facto reward signal for post-training reasoning agents. On long-horizon, multi-step tool-use tasks these judges exhibit silent reward hacking: they systematically overrate verbose, well-formatted, or confidently phrased trajectories while underrating substantively correct but terser solutions. The aggregate pairwise accuracy that existing benchmarks report conflates competence with surface conformity. We introduce RubricGuard, a controlled paired-probe auditing protocol that decomposes judge behavior on long-horizon trajectories into six single-target failure modes — length, format, confidence, citation-fabrication acceptance, intermediate-step credit forgery, and distractor susceptibility — each instantiated by probe pairs that hold task semantics fixed while toggling one confound against a pre-committed falsification tolerance. Three diagnostic metrics — Failure-Mode Coverage, Inversion Rate, and Multi-Step Credit Fidelity — turn aggregate accuracy into an axis-level reliability profile: across six recent judges and four long-horizon benchmarks, aggregate accuracy between 76% and 84% coexists with per-axis inversion rates from 5.8% to 34.1%, and every judge is falsified on at least one axis at the 0.10 tolerance. A training-free inference-time intervention, Rubric-Constrained Decoding (RCD), conditions the judge on a per-task rubric and a contrastive sibling pair, recovering 51%–68% of inversions at a 1.4× latency overhead. The audit harness, a probe subset, and the decoder ship with the submission; the full 28,800-probe set will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.