CHRF-SAM3: Constrained Heterogeneous Representation Fusion for Cross-Domain Few-Shot Adaptation
Abstract
Cross-domain few-shot visual learning aims to adapt models to unseen visual domains with only a small number of labeled samples. Benefiting from large-scale pretraining, SAM3, as a new-generation vision foundation model, exhibits strong zero-shot generalization, yet it struggles to effectively adapt to novel cross-domain visual distributions under extremely limited supervision. To address this issue, we introduce DINOv3, which is heterogeneous from SAM3 and yields multi-scale visual representations learned via a distinct pretraining paradigm. Specifically, we adopt its convolutional variant. Building upon this, we develop Convolutional Residual Fusion (CRF) and Layer-Shared Conditional LoRA (LSC-LoRA) to inject the heterogeneous multi-scale representations into the visual encoder and enable conditional low-rank adaptation. On the decoder side, we employ SoMA to constrain low-rank parameter updates. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.