acceptodds
Under review as a conference paper at ICLR 2027

Beyond Endpoints: Bézier Paths for Visual Prompt Tuning

Abstract

Visual prompt tuning (VPT) adapts a frozen vision foundation model by learning a few tokens shared across all images. When prompts are updated as data shift or tasks change, the distribution of image features moves from one endpoint to another. Prior theory asks which endpoints prompts can reach; we ask whether they can travel between them along the shortest path. Because shared prompts compete with each image's own content for attention, they often cannot. We quantify this restriction as the extra transport cost that shared attention imposes over the shortest path. In an exactly solvable attention model, the reachable features form a cone; unrolling it into a plane gives the exact minimum action for every pair of endpoints and explicit optimal prompts on short arcs. The analysis reveals a sharp threshold in content normalization, which measures how much attention each image's own tokens absorb: with at most two distinct levels, prompts can follow the shortest path, whereas three or more force a strictly positive extra cost unless the endpoints differ only by rescaling. This gap persists with more prompts, small content perturbations, and aligned multihead outputs. On ViT-B/16 with VTAB-1k, a quadratic Bézier prompt path that minimizes feature action reduces held-out transport action relative to linear and cross-entropy-trained paths between the same endpoints, while cross-entropy-trained paths retain higher minimum accuracy. Reachable endpoints do not guarantee a reachable shortest path: how prompts move matters, not only where they end.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.