acceptodds
Under review as a conference paper at ICLR 2027

Learning What to Reject: Counterfactual Preference Alignment for Emotion–Cause Pair Extraction

Abstract

Emotion–cause pair extraction (ECPE) identifies emotion clauses together with the clauses that explain them. A persistent challenge is that ECPE benchmarks are strongly locality-skewed: annotated causes often occur near their emotions, making relative position a highly predictive shortcut for cause selection. Models can therefore favor nearby plausible clauses without learning the distinction that determines the annotation-supported cause. Existing methods mainly strengthen candidate representation or reasoning, but standard positive supervision provides little targeted contrast against shortcut-compatible alternatives. We introduce CPA-DPO, a counterfactual preference-alignment framework that turns such alternatives into explicit competitors. CPA-DPO constructs reliable, challenging, and controlled preferences: verifier-defined eligibility filters counterfactuals, hardest-valid selection chooses the most competitive eligible alternative, and position-balanced formation prevents position from deterministically revealing the preference label. Standard DPO then optimizes these task-specific comparisons. On the matched Qwen2.5-32B setting, CPA-DPO improves ECPE-CN pair F1 from 83.62 to 85.43 over Vanilla-DPO (). Its advantage grows to F1 on LONG pairs and on Rebalanced-CN, while independent hard-negative accuracy rises from 87.92 to 91.35.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.