acceptodds
Under review as a conference paper at ICLR 2027

LEARNING HIERARCHICAL FUSION OF INTERMEDIATE REPRESENTATIONS ACROSS VISION FOUNDATION MODELS

Abstract

Modern vision foundation models expose rich hierarchies of intermediate rep- resentations, yet downstream transfer commonly reduces each encoder to its final feature; this discards potentially useful information across depth within an encoder and across independently trained encoders. We introduce Hierarchical Fusion At- tention (HIFA), a hierarchical attention mechanism that first aggregates interme- diate representations within each frozen vision backbone and then fuses the result- ing summaries across backbones, operating on compact global layer descriptors rather than dense spatial maps so as to separate the selection of useful representa- tion levels from the combination of heterogeneous encoders. On VTAB-1k, HIFA achieves substantially higher mean accuracy than concatenation of final-layer rep- resentations, with particularly large gains on Structured tasks: conventional con- cat reaches 0.660, a packed-final concat control 0.671, final-only HIFA 0.702, an architecture-matched final-token-only HIFA 0.704, and full HIFA 0.719. The near-equality of final-only and matched final-token indicates that the remaining numerical lift to full HIFA is associated with access to intermediate depth under the same hierarchical architecture; this depth-related difference is numerical and task-dependent rather than a statistically established uniform cross-task improve- ment. Beyond predictive performance, intra-backbone attention reveals alloca- tion across network depth and inter-backbone attention reveals encoder contribu- tions; combined with test-time masking interventions, these analyses show that the trained fusion mechanism depends strongly on specific depths and encoder summaries.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.