Hierarchical Ensemble Adapter Tuning for Few-Shot Adaptation
Abstract
Parameter-efficient few-shot adaptation usually trains one lightweight adapter of fixed bottleneck rank. That single rank commits adaptation to one hypothesis, which often lacks the robustness required for few-shot scenarios. In this paper, we propose Hierarchical Ensemble Adapter Tuning (HEAT), a framework that trains an ensemble of lightweight adapters with varying bottleneck ranks, each imposing a different structural inductive bias. Based on multiple adaptation hypotheses, HEAT expands the model's representational capacity by learning nested, rank-constrained features: a shared core of dominant task directions that each higher-rank adapter refines. To leverage this hierarchical structure at inference time, we introduce an uncertainty-aware aggregation strategy that dynamically weights each adapter's output according to its instance-wise predictive entropy. Analysis of the learned representations reveals that the adapters form nested subspaces where each of them sharing a common core subspace and that entropy weighting systematically exploits this structure, amplifying the most confident adapter on a per-sample basis. Across eleven few-shot benchmarks, HEAT achieves state-of-the-art accuracy all within the same parameter budget as a single CLIP-Adapter.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.