acceptodds
Under review as a conference paper at ICLR 2027

CSA-RIT: Learning Noise-Resilient Multimodal Recommendation via Coarse Semantic Alignment and Robust Interest Tree

Abstract

Graph-based multimodal recommenders typically build an item-item graph per modality and aggregate neighbor representations by graph convolution. Two things decide the outcome of this aggregation. The first is what the modality features being averaged express, and the second is which items enter a neighborhood. Both are settled before the first propagation step, and repeated smoothing spreads any error into every representation. On the feature side, an item's visual and textual features are typically extracted independently and often inconsistent. We call this content noise. On the graph side, existing methods often introduce noisy links when building item-item graphs. We call this structural noise. We propose CSA-RIT, a noise-resilient framework that treats content and structural noise as two sources of a single pre-propagation error and repairs both. Coarse Semantic Alignment (CSA) places a coarse discrete bottleneck ahead of graph propagation and aligns the same item across modalities only in a shared coarse code, so that the signals to be averaged share a common semantic basis. Robust Interest Tree (RIT) augments each frozen modality graph with a short, order-aware co-occurrence tree, supplying a stable prior that resists this structural noise. On three standard MMRec benchmarks spanning catalog sizes from 7k to 63k items, CSA-RIT is best on every metric, improving over the strongest baseline by 8.9% on average and by 13.7% on the largest split. Ablations confirm complementary gains, and RIT keeps its recall under injected structural noise. Our implementation is available at https://github.com/paper-ai212/CSA-RIT.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.