acceptodds
Under review as a conference paper at ICLR 2027

Your Model Already Knows, Don't Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models

Abstract

Few-shot adaptation of vision-language models (VLMs) is difficult for specialized aerial, industrial, medical, and other imagery when only ten annotated images are available. Discrete prompt optimizers such as GEPA and DetPO keep the model frozen but search text without numerical gradients; low-rank adapters use gradient descent but learn millions of task-specific parameters. We revisit soft prompting, which optimizes a few continuous input tokens by gradient descent against a frozen backbone, and find that two defaults inherited from NLP hold it back: prefix placement and semantic initialization. Placing the tokens at the cross-modal boundary, between image and text, and initializing them from a semantically empty space token unlock its performance. On Roboflow20-VL, the resulting prompt uses 7,168 trainable parameters on average, outperforms GEPA and DetPO under a shared multi-class protocol, and matches the best LORA configuration at 14.2 mAP@50:95 with roughly 24,000× fewer trainable parameters. The learned tokens behave like prompts rather than weights: they transfer to a newer model version without retraining, and SOFT2HARD verbalizes them into an editable instruction that matches DetPO. The placement principle also carries beyond detection: on RoboCasa, a prompt that reaches the action expert of a frozen robot policy raises two near-floor tasks from 5.0% to 23.3% and 31.7% success. Our results suggest that modern VLMs already encode much of what specialized domains require; the task is not to teach them, but to learn how to ask.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.