Aligning within Pair-Specific Subspaces for Noisy Correspondence Learning
Abstract
Noisy correspondence introduces semantically mismatched image-text pairs as positives, inducing erroneous attraction while degrading the supervision available from correct correspondences. Existing methods mitigate this problem through reliability estimation, robust optimization, and correspondence correction. However, pair-level weighting mainly controls the learning strength of pairs with different confidence levels, without directly identifying suitable directions for auxiliary alignment. Auxiliary constraints based on similarities between samples may also propagate correspondence errors to other pairs and interfere with learning reliable correspondences. We propose APS (Aligning within Pair-Specific Subspaces), which extracts multiview consistency evidence from each observed pair to determine the space for auxiliary alignment. Specifically, a frozen teacher model, obtained through reliability-weighted warm-up, evaluates the consistency of image-text view combinations. Their selected joint representations are then used to construct a pair-specific subspace. Correspondence confidence jointly weights the global contrastive objective and the auxiliary subspace objective, strengthening image-text consistency for higher-confidence pairs while suppressing unreliable auxiliary constraints. Experiments on Flickr30K, MS-COCO, and CC120K demonstrate the effectiveness and robustness of APS under synthetic and real-world correspondence noise.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.