Selective Tie Learning: Mitigating Spurious Correlations in Preference Optimization
Abstract
Preference optimization can learn spurious correlations, and augmenting the preference data with ties (equally preferred responses) can reduce spurious reliance. However, ties differ in their corrective effect, making their selection important. For log-linear policies, we derive each tie's influence on spurious features and show that taking the largest predicted reductions is first-order optimal among subsets of equal size. Computing these influences requires knowledge of the spurious features and the curvature of the training objective, which is often unavailable or expensive to estimate. We therefore study the absolute implicit reward margin as a proxy for tie influence, giving a selection rule that requires two forward passes per tie. To test the rules beyond the theoretical setting, we develop practical influence estimates using a counterfactual proxy for reliance on spurious features and a generalized Gauss–Newton approximation to curvature. Across log-linear policies, neural models, and large language models, our selection rules consistently achieve comparable reductions in spurious reliance with fewer ties than uniform selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.