acceptodds
Under review as a conference paper at ICLR 2027

From Divergence Scores to Auditable Evidence: Exact-TV Anchors for Reward-Interface Dependence

Abstract

Trustworthy paired-rollout audits require an observable behavioral target and a conclusion that survives independent administration. We meet both requirements for reward-aware agents by measuring how a frozen policy's action distribution responds to task-preserving edits of shaping meters, verifier messages, judge feedback, and reward-model outputs. The Paired Interface Response Audit (PIRA) admits edits through ground-truth preservation tests, estimates an intervention-indexed response profile from matched rollouts, and returns a finite-sample rank-conformal Holm decision (PIRA-Holm) together with a learned ranking score (PIRA-L). Three open-weight exact-TV anchors expose the target action-distribution shift directly; frozen reruns and low-overlap external banks then test execution and bank-design invariance. PIRA-Holm attains AUROC and PIRA-L attains on -bench-Exact, SWE-Gym-Exact, and MLE-Bench-Exact. Independent operators reproduce these results within AUROC. Three full external handoffs with – semantic overlap reach , , and , retain worst-family AUROC of at least , and keep Benign-Hard false-positive rates below . These results turn reward-interface sensitivity from a score correlate into an exact-anchored interventional measurement and give paired-rollout audits a reproducible evidentiary bar.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.