acceptodds
Under review as a conference paper at ICLR 2027

Reading the Velocity Field: Closed-Form Analysis of a Cross-Modal Flow Adapter

Abstract

Flow-matching adapters specialise a vision–language model from a few labelled images: they regress a velocity field carrying image features toward the text embedding of their class, then integrate it at test time. We show the population optimum of this objective is available in closed form. Inverting the interpolant removes the image feature from the regression target, and in the noise-free setting these methods use, what remains is geometry: the field's value is a convex combination of the text anchors, the Euler iterate has an exact form, and the classifier logits split into a zero-shot term and a posterior smoothed by the anchor Gram matrix . That split is useful. Reading the posterior off a trained field and integrating the field are the same estimator applying and , so neither dominates and the better choice depends on how well conditioned is. Interpolating them recovers both as limits, and the weight is not free: since the residual covariance in anchor coordinates is , the forward map itself, the interpolating estimator is Bayes-optimal with . This yields a decoder with no tuned hyperparameters needing one forward pass and no ODE solve. On the five difficult benchmarks of jiang2026exploring it modestly but consistently improves on integration with almost tuning free. A matched-form control attributes the gain to the field rather than the estimator: applied to raw features the same decoder provides at par performance over zero-shot. Our results highlight that principled closed-form analysis can simplify inference while modestly improving few-shot adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.