acceptodds
Under review as a conference paper at ICLR 2027

When Can Counterfactual Audits Certify a Fine-Tuning Gain?

Abstract

Counterfactual post-training expands supervision through generated edits, but neither more edits nor greater predictive consistency establishes a gain over factual training. Certifying that gain requires an independent comparison with enough power to detect it under the declared nuisance shifts. We study this evidence requirement through paired losses: common model responses can cancel, while additional edits cannot remove uncertainty across task anchors. Shared-anchor certificates and finite-sample power bounds distinguish these two sources of cost. An independent pilot–plan–certify procedure then selects the audit size, sibling count, estimator, and confidence allocation from a frozen menu, accounting for uncertainty in the planning moments themselves. Its guarantees separate false certification, pilot reliability, and final power conditional on a feasible plan. Controlled experiments isolate the savings from pairing and anchor reuse; across three planning laws, 109 of 144 pilots return feasible plans, each successful in all 256 fresh replications, but their budgets remain conservative. A prospective Qwen/ReCoRD audit illustrates covered-shift certification with a different, finite-catalog design. The practical lesson is to budget for uncertainty in the audit plan, not just in the final result: a valid certificate and an economical, adequately powered audit are distinct achievements.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.