Verify What You Can Parse:Prompt-Gated Residual Reranking for Reward Models
Abstract
Reward models must assess both semantic quality and literal instruction compliance, although constraints such as exact counts, required phrases, and output formats are often more reliably evaluated by explicit programs. We introduce SpecGate, a post-hoc reranking layer for frozen reward models. A prompt parser restricts intervention to supported constraints, an executable checker measures literal compliance, and a -parameter ensemble predicts complementary soft compliance signals. A spread gate suppresses learned corrections when their evidence is weak. Using a configuration fixed before the reported replay, SpecGate improves official RewardBench 2 scores from to on Llama-3.1-8B and from to on Qwen3-8B. The aggregate changes are confined to Precise Instruction Following, where the method produces and net corrections, respectively. These corrections satisfy , matching the benchmark’s six-domain macro-average. Paired bootstrap confidence intervals are and percentage points, with exact McNemar . Ablations show that the executable checker contributes most of the improvement, while the learned component provides a smaller, model-dependent benefit. These results support selective verification for candidate reranking when prompts expose mechanically checkable requirements; they do not imply a general replacement for semantic reward modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.