Disagreement Is Not Degradation: Evaluating Reliability in Multimodal Sentiment Analysis
Abstract
Multimodal reliability evaluations can conflate annotation disagreement, input degradation, task utility, and output-format compliance. We audit these distinctions with three omni-modal models: a 21,714-request, seven-view evaluation on 1,034 CH-SIMS v2 clips, followed by paired raw-media transformations and fixed MELD and CMU-MOSEI cohorts. Adding a transcript to audio-visual input increases Qwen2.5-Omni's error on the annotation-disagreeing CH-SIMS subset, while the corresponding MiniCPM-o estimate remains uncertain. A training-median control reveals different target distributions: none of the three model-minus-prior intervals on the disagreeing subset excludes zero. Qwen3-Omni's format failures further separate schema compliance from task performance. Narrow syntax rules developed post hoc on the original run recover its malformed numerical responses; fixed before the follow-up, complete-fence normalization also restores high coverage on the external tasks. Additional modalities do not yield uniform gains: on the locked MELD cohort, MiniCPM-o's text accuracy exceeds its text-audio-visual accuracy by 9.96 percentage points despite complete strict coverage. Paired physical transformations have nonuniform predictive effects, and an earlier matched-supervision feature study does not support a factorized reliability gate. The evidence supports joint reporting of absolute error, paired input effects, train-only prior controls, output coverage, and distinct failure causes. It does not establish a human-validated perceptual-quality effect or a superior fusion architecture.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.