acceptodds
Under review as a conference paper at ICLR 2027

ACA: COUPLING RELATIONAL ALIGNMENT AND COMPLEMENTARY REGULARIZATION FOR VISION ENCODER FUSION

Abstract

Combining multiple vision foundation models has become an important strategy for improving visual representations. A common approach concatenates their [CLS] tokens and trains a task-specific predictor, but does not explicitly constrain patch-level relations. Inspired by DINOv3’s Gram anchoring, which preserves patch relations during self-supervised training, we study relational alignment across encoders. Uniform Gram matching penalizes every pairwise discrepancy, whereas fusion calls for aligning selected relations while retaining differences between representations. We introduce Alignment-guided Complementary Aggregation (ACA), which couples selective relational alignment with regularization of the complementary residual. Global features provide importance weights, while projected patch features determine a shared gate for weighted alignment, a mean-gate penalty, and a variance reward on the weighted complementary residual. Both encoders are jointly trained and retained for prediction. Across four classification settings on COCO and ImageNet-1K, ACA improves over matched-pair concatenation by 0.25–0.94 percentage points and outperforms uniform Gram alignment in three. On ImageNet-Hard, the full objective reaches 69.98% and exceeds all six one- and two-term variants. ACA also exceeds direct summation of the importance, residual, and gate signals by 1.01 points on ImageNet-Hard and 0.67 points on COCO. These results support coupling relational alignment with complementary residual regularization for encoder fusion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.