LEARNING HIERARCHICAL FUSION OF INTERMEDIATE REPRESENTATIONS ACROSS VISION FOUNDATION MODELS
Abstract
Modern vision foundation models expose rich hierarchies of intermediate rep- resentations, yet downstream transfer commonly reduces each encoder to its final feature; this discards potentially useful information across depth within an encoder and across independently trained encoders. We introduce Hierarchical Fusion At- tention (HIFA), a hierarchical attention mechanism that first aggregates interme- diate representations within each frozen vision backbone and then fuses the result- ing summaries across backbones, operating on compact global layer descriptors rather than dense spatial maps so as to separate the selection of useful representa- tion levels from the combination of heterogeneous encoders. On VTAB-1k, HIFA achieves substantially higher mean accuracy than concatenation of final-layer rep- resentations, with particularly large gains on Structured tasks: conventional con- cat reaches 0.660, a packed-final concat control 0.671, final-only HIFA 0.702, an architecture-matched final-token-only HIFA 0.704, and full HIFA 0.719. The near-equality of final-only and matched final-token indicates that the remaining numerical lift to full HIFA is associated with access to intermediate depth under the same hierarchical architecture; this depth-related difference is numerical and task-dependent rather than a statistically established uniform cross-task improve- ment. Beyond predictive performance, intra-backbone attention reveals alloca- tion across network depth and inter-backbone attention reveals encoder contribu- tions; combined with test-time masking interventions, these analyses show that the trained fusion mechanism depends strongly on specific depths and encoder summaries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.