Beyond Person-Cue Suppression: Three-View Matched Contrastive Correction for Visual Cultural Attribution
Abstract
Multimodal large language models (MLLMs) can assign different cultural origins to the same content when depicted with different people. However, suppressing person-related cues alone does not ensure correct attribution: it may obscure relevant cultural evidence while leaving attribution errors unresolved. To address this limitation, we propose Three-View Matched Contrastive Correction (TVMCC), which uses the visual intervention itself to construct a reference for correcting the suppressed prediction. Task-aware masking produces a person-suppressed view and a complementary person-focused view from the same image while seeking to preserve the queried content. By contrasting candidate preferences in the original and person-focused views, TVMCC revises the suppressed ranking using a reference matched to the current image and intervention, without additional training or test labels. Across six MLLMs on Cultural Counterfactuals and MixCuBe, TVMCC outperforms original-image predictions in all 18 model–task settings, achieving the highest reported accuracy in 17. Nationality attribution gains on Cultural Counterfactuals range from 10.4 to 16.9 percentage points, exceeding person suppression on all six models. Reference shuffling on Cultural Counterfactuals reduces accuracy even with matched correction magnitudes, supporting the contribution of same-image correspondence beyond correction strength.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.