Modality Disagreement Guided Fine-Tuning for MLLM-Based Sequential Recommendations
Abstract
Multimodal sequential recommendation combines interaction histories with item images and text to model users' preferences. Recent approaches fine-tune multimodal large language models (MLLMs) to rank candidate items from these inputs. However, candidates that are similar in one modality may differ substantially in the other. We ask whether such cross-modal disagreement can serve as a content-based criterion for deciding which comparisons to emphasize during recommendation training. We propose MoDiFy-Rec, a modality-disagreement-guided fine-tuning method that uses these differences to weight candidate comparisons in MLLM-based sequential recommendation. Specifically, MoDiFy-Rec measures disagreement using frozen visual and textual representations. It first maps pairwise similarities to percentiles in their respective training-set distributions, making their relative positions comparable across modalities. An exponential kernel and symmetric normalization then convert the absolute differences between these normalized similarities into fixed positive comparison weights. Following cross-entropy initialization, MoDiFy-Rec fine-tunes the MLLM with full-candidate cross-entropy and a weighted pairwise logistic loss comparing the observed next item against every other candidate. For fixed contexts with positive conditional label probabilities, we show that the combined objective preserves cross-entropy’s population minimizers over unrestricted scores for any nonnegative auxiliary coefficient. On three recommendation datasets, MoDiFy-Rec improves ranking quality over cross-entropy fine-tuning and pairwise baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.