Beyond Text Anchors: Dual-Anchor Flow Matching for Few-Shot and Incremental Adaptation
Abstract
Flow-based adaptations improve few-shot recognition with vision–language models by transporting visual features toward their corresponding class textual prototype. However, they implicitly assume that a semantically correct textual prototype is a reliable and discriminative transport endpoint. This assumption may fail when textual class geometry is crowded or misaligned with the visual distinctions required by a downstream task. A support-derived visual prototype offers a natural alternative by capturing task-specific appearance cues, but can itself be unreliable under scarce supervision. This complementary reliability motivates **DualFlow**, a dual-anchor flow-matching framework that retains textual and visual prototypes as two endpoint hypotheses rather than committing solely to either one. DualFlow uses a shared flow backbone with two anchor-specific heads to learn textual and visual-guided velocities, which are adaptively coordinated by a support-derived reliability coefficient shared with the classifier anchors. Beyond conventional few-shot adaptation, this dual-anchor transport-based paradigm is naturally suited to few-shot class-incremental learning (FSCIL): newly arriving classes can be incorporated by appending their textual and visual anchors and adapting only the lightweight velocity fields, while frozen vision–language encoders preserve a shared representation geometry for previously learned classes. Experiments across conventional and incremental few-shot adaptation benchmarks show that DualFlow substantially improves over text-only flow adaptation and complements parameter-efficient VLM adaptation, highlighting endpoint reliability as a critical factor in discriminative feature transport.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.