Volterra Flow Matching with Memory for Few-Shot classification
Abstract
While pre-trained vision-language models provide a coarse alignment between image and text embeddings, they often require parameter-efficient fine-tuning (PEFT) to adapt to downstream tasks like few-shot learning. Existing PEFT methods, including prompt tuning, LoRA, and adapter-based approaches, typically perform one-step updates, while recent multi-step methods model refinement via learned velocity fields. However, these approaches typically condition the dynamics on the current state, leaving the potential predictive role of trajectory history underexplored. We hypothesize that, especially in low-data regimes, the instantaneous feature state can be insufficient for predicting the appropriate alignment dynamics. To capture this structure, we propose a Volterra Flow Matching (VFM) framework that introduces a memory variable to encode the history of feature evolution. This memory is constructed via a Volterra integral with an exponentially decaying kernel, leading to a coupled dynamical system that governs both the current state and its accumulated history. Practically, we construct interpolation trajectories between image and text features, and train a lightweight neural operator to predict the drift conditioned on both the instantaneous features and their memory state. During inference, the learned dynamics are integrated using a small number of steps, enabling efficient multi-step refinement. Extensive experiments on standard few-shot benchmarks demonstrate modest yet consistent gains over existing PEFT and Flow Matching Alignment methods, while controlled prediction-risk probes show that the memory state contains additional velocity-predictive information.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.