acceptodds
Under review as a conference paper at ICLR 2027

MGSPRec: Multimodal Sequential Recommendation with Multi-Granularity Shared and Private Representation Learning

Abstract

Multimodal sequential recommendation uses visual and textual content to enrich the modeling of users' dynamic preferences. However, raw modality features may contain information that is irrelevant to recommendation. Meanwhile, different modalities have distinct representation spaces and semantic structures. Direct parameter sharing or forced alignment across modalities may therefore introduce cross modal interference and weaken valuable modality specific semantics. To address these issues, we propose MGSPRec, a multimodal sequential recommendation method based on multi granularity shared and private representation learning. MGSPRec first uses behavior guided signals to refine visual and textual representations. This process suppresses recommendation irrelevant noise while preserving modality specific semantics. It then employs separate Transformer encoders for ID, visual, and textual sequences to reduce interference across modalities. Furthermore, ID based collaborative information is used as a semantic anchor. Shared and modality specific representations are learned at both the item and sequence levels. Experiments on three real world datasets, Baby, Pantry, and Office, show that MGSPRec consistently outperforms representative baselines. Compared with SGP4SR, MGSPRec achieves average relative improvements of 4.80%, 3.69%, and 2.29% over HR and NDCG at different cutoff values on Baby, Pantry, and Office, respectively. Ablation studies and parameter sensitivity analyses further validate the effectiveness of the proposed components.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.