acceptodds
Under review as a conference paper at ICLR 2027

Candidate Completion for Exact Off-Policy Calibration of Best-of-M Policies

Abstract

Choosing the highest-scoring action from several candidates changes which prediction errors a controller encounters. A bound calibrated on the original policy's data can therefore lose its intended coverage. We introduce candidate completion, which calibrates best-of-M policies using logged trajectories and samples from the preserved data-collection policy. It places each logged action among fresh competitors and tests whether the deployed rule selects it. Applying these tests along a trajectory yields exact target-policy samples without evaluating action densities or collecting new environment outcomes. Ordinary and weighted calibration then provide finite-sample prediction bounds. In repeated-selection experiments on directly observed control rewards, completion corrects Walker2d-high coverage from 78.7% to 90.4% under two-step best-of-four selection. Across all 27 setting/group comparisons, its mean absolute coverage difference from a target-data oracle is 0.69 percentage points, versus 2.12 for behavior calibration. In controlled equal-budget comparisons with 500 logged actions, completion achieves 90.4–91.7% coverage for 4–16 candidates, while score-CDF weighting achieves 86.6–87.3%. Additional selection sweeps, tied-score experiments, and complete-trajectory checks support the sampling and calibration guarantees.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.