acceptodds
Under review as a conference paper at ICLR 2027

Detecting Structured Effects in Preference Data: A Fisher Information Perspective

Abstract

Preference comparisons selected to improve reward-model fit need not be informative for detecting interactions among evaluation dimensions. We study this selection problem through nuisance-adjusted Fisher information, which combines preference variance with the interaction signal remaining after accounting for main effects. This criterion links labeling budget to a local detection boundary and motivates Cond-Opt, a frozen-pilot selector accompanied by checks of selected-set identifiability and information retention. In fixed HelpSteer2+3 and UltraFeedback pools, the top 1% of comparisons account for 92.6% and 89.8% of estimated interaction information. In matched-budget simulations with split-null calibration, Cond-Opt achieves higher power than unresidualized interaction selection and one-shot full-parameter Fisher leverage across all six evaluated budget–noise settings at γ* = 0.5. Against uniform sampling, its median gain is 58.1 percentage points within the prespecified uniform-power window of 20–80% in a separate grid. On HelpSteer2+3 features with simulated labels, it reduces the 50%-power detection boundary by 36%. Tests on observed human labels do not confirm the specified interaction at the evaluated scale. Together, these results identify the value of targeting residual interaction information and provide a framework for assessing which comparisons can inform a specified preference question.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.