acceptodds
Under review as a conference paper at ICLR 2027

PIVOT: Risk-Controlled Routing of Verifier Feedback for Reliable Policy Improvement

Abstract

Post-training relies on verifiers with different reliability, coverage, and cost. Existing routers principally estimate label accuracy, although learning depends on the update induced by feedback. We introduce PIVOT, a sequential training-time router that treats verifier feedback as a policy intervention. It predicts reliability adjusted directional alignment before acquisition, maps heterogeneous outputs to preference graphs, and gates the resulting optimizer-native update. Across Llama-3.1-8B-Instruct and Qwen2.5-7B-Instruct in four domains under a common cost cap, PIVOT reaches 68.4 trusted utility. It improves on common-pool BayesianRouter-H by 1.6 points and uncertainty routing by 2.3 points, with worst slice utility of 61.1. At the pre-specified α = 0.05 budget, held-out gold-graph inconsistency (GG-risk) is 0.047 at 82% coverage. Setup cost amortizes after 37.6K prompts relative to BayesianRouter-H and 67.6K relative to uncertainty routing. Fully stale reuse after policy shift raises GG-risk from 0.048 to 0.082, delimiting the phase-wise guarantee.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.