acceptodds
Under review as a conference paper at ICLR 2027

Not All Preferences Are Strict: Tie-Aware Preference Optimization for Multi-Image MLLMs

Abstract

Preference optimization methods for multimodal large language models (MLLMs) universally assume strict pairwise preferences. We show that this assumption sys- tematically fails in multi-image settings, where responses frequently tie—arriving at the same correct conclusion through different image subsets, visual cues, or levels of detail. We formalize a taxonomy of four structurally distinct tie categories— detail, style/framing, evidence-path, and order-invariance ties—among which evidence-path ties are unique to multi-image reasoning and order-invariance ties predominantly arise in multi-image contexts. Forcing such ties into binary labels introduces spurious preference boundaries that manifest as verbosity bias, position bias, and evidence-path overcommitment. We propose TieDPO, a tie-aware pref- erence optimization objective that preserves standard DPO on strict pairs while regularizing tie pairs toward indifference via a margin-based penalty. To support training, we construct MultiTie-22k, a multi-image preference dataset with ternary labels and structured tie-category annotations. We further introduce TieBench-MI, a diagnostic benchmark of 2400 pairs for evaluating non-strict preference handling. Experiments show that TieDPO matches or exceeds DPO baselines on standard multi-image and single-image tasks while significantly improving tie recognition and reducing preference biases.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.