GAFT: Geometry-Aware Fine-Tuning for Foundation Models
Abstract
Adapting Vision Foundation Models (VFMs) in low-label regimes remains challenging despite their strong transferable representations. While semi-supervised learning (SSL) can mitigate label scarcity, most existing approaches treat unlabeled data primarily as noisy supervision, without explicitly accounting for the geometry of the learned feature space. In this work, we develop a theoretical framework showing that unlabeled data improves generalization by revealing the structure and density of the representation space, yielding a model-adaptive and data-adaptive bound in which labeled samples act as geometric anchors and local prediction robustness governs expected risk. We then propose GAFT (Geometry-Aware Fine-Tuning) which introduces geometry-aware regularization terms that integrate seamlessly into existing SSL pipelines. Extensive experiments on the Visual Task Adaptation Benchmark (VTAB), across multiple fine-tuning strategies, demonstrate that GAFT consistently improves both accuracy and robustness over strong baselines, highlighting the effectiveness of geometry-aware SSL for VFMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.