acceptodds
Under review as a conference paper at ICLR 2027

Synergy-Aware Multimodal Disentanglement Learning

Abstract

Disentangling a multimodal representation into factors for different information types makes it inspectable and reusable: the evidence each factor carries can be probed or recombined. Partial Information Decomposition (PID) provides natural coordinates for this, dividing task information into redundant, unique, and synergistic parts. However, to our knowledge, existing self-supervised multimodal disentanglement stops at a shared/private split, and interaction-oriented methods capture synergy only inside one fused embedding. We propose **URSyn** (**U**nique, **R**edundant, **Syn**ergistic), a self-supervised framework that learns separately accessible factors for all three types: unique factors are repelled from redundant ones, redundant factors are supported by every single modality, and the synergy factor is anchored to the full input while repelled from all single-modality references. We show that idealized versions of these constraints bound the task information of the redundant and synergy factors by redundancy- and synergy-related information quantities, without requiring PID values to be estimated. On a controlled benchmark, each factor best predicts its designated target, and the synergy factor *alone* reaches 83.5% on a joint-only task, exceeding the complete representations of CoMM (70.4%) and InfMasking (77.0%). On six real-world benchmarks, **URSyn** matches or exceeds our CoMM reproduction in mean performance under the same protocol and is close to InfMasking's reported results.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.