Prototype-Guided Recalibration and Inter-Modal Signal Reshaping for Multi-Modal Sequential Recommendation
Abstract
Multi-modal sequential recommendation incorporates textual and visual features to alleviate interaction sparsity. However, while uncalibrated multi-modal integration may improve tail-item recommendations, it often degrades head-item performance by distorting their robust collaborative representations. To address this limitation, we propose Prototype-Guided Recalibration and Inter-Modal Signal Reshaping for Multi-Modal Sequential Recommendation (PRISM), which recalibrates multi-modal features before ID-based sequence modeling. Cross-Modal Semantic Reshaping separates shared semantics from modality-specific cues, aligning the former while using prototypes to preserve recurring patterns in the latter. Frequency-Gated Sequence Recalibration jointly analyzes textual and visual amplitude spectra to adaptively regulate each modality's sequence variations without uniformly smoothing informative dynamics. Cross-attention then integrates the recalibrated sequences with the ID-based sequence for prediction. Experiments on four datasets show that PRISM outperforms representative baselines and improves performance on both head and tail items.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.