From Prediction to Selection: Outcome Supervision for Closed-Loop Best-of-
Abstract
Accurate outcome prediction across states need not imply accurate action ranking within a state. We study this distinction in repeated best-of-five control with a frozen actor. Alongside a theoretical construction, we compare four critic-training conditions: Base, ordinary-rollout augmentation, samestate counterfactual (CF) outcome supervision, and a control that replaces candidate targets with their within-state mean. Across 40 LIBERO tasks, CF yields absolute success-rate gains over Base of 4.44% and 3.75% for IQL and CQL, respectively, with confidence intervals across tasks and initializations spanning zero. In the PointMaze UMaze Fast64 configuration, CF improves normalized return by 5.20 points over Base and 2.89 points over the meantarget control, with positive confidence intervals for both gains. Candidatespecific gains remain uncertain in LIBERO, and CF effects vary across UMaze configurations. Local LIBERO ranking differences have confidence intervals spanning zero. The study separates the overall effect of outcome supervision from the effect of retaining candidate-specific targets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.