A Modality with No New Information Can Improve Multimodal Fusion
Abstract
A fusion model that beats its strongest unimodal component is often credited with using information from the added modality. Adding a branch, however, also changes parameterization and optimization. We propose a rotation-placebo audit that separates the two: beside the real modality, retrain the same fusion architecture with an invertible orthogonal rotation of the raw base input, which adds zero Shannon information, and with parameter-matched copy and Gaussian arms. A gain that the placebo reproduces cannot be attributed to new information. In the knowledge-tracing deployment that motivated this work, the audit overturns the original attribution: the placebo raises area under the curve by an amount comparable to the gain that had justified collecting student feedback, including records from minors. It also raises it in the public Computer Science Education Data Mining (CSEDM) challenge, though there its log-loss gain (+0.005 bits) is below our 0.01-bit floor. The audit is also specific. It stays near zero where a real modality helps substantially: on frozen Contrastive Language–Image Pretraining (CLIP) features, in a registered audit of the official Adaptive Gradient Modulation implementation, and in a second public audit that reproduces a +10.81-point real-modality gain while copy and rotation gain only +1.96 and +0.65 points. Where the placebo fires, registered interventions control it. It reduces held-out cross-entropy in all 12 registered raw-input audio–visual MNIST (AV-MNIST) configurations (60 paired runs). Saturating the task removes it, and rank-256 principal-component truncation and whitening make all six settings fail our prespecified positive-placebo criterion while preserving a positive real-modality gain (largest residual +0.022 bits, positive in 4/5 seeds). A converged probe adds a one-sided certificate that bounds the cost of ignoring a candidate, with no violations among 42 eligible cases. A base-versus-fusion comparison alone cannot attribute a gain to new information; the audit makes that attribution testable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.