VAPOR: Automatic Prompt Optimization for Vision-Language Models via Visual Reflection
Abstract
Automatic prompt optimization (APO) has emerged as a gradient-free alternative to fine-tuning for adapting vision-language models (VLMs). However, existing APO methods inherit a modality mismatch: they optimize visual decision-making through language-only feedback. This mismatch is especially limiting in few-shot settings with fuzzy visual boundaries, where failures often arise not from missing task instructions but from misidentified visual evidence. We propose , an automatic prompt optimization framework that closes this feedback gap through visual reflection. VAPOR compares per-example reasoning trajectories between erroneous and corrected predictions, distills their differences into atomic visual decision rules, and records negative constraints that suppress recurring spurious rationales. We treat model-generated rationales as candidate diagnostic signals rather than faithful explanations, and admit a rule into the optimized prompt only after behavioral verification. Rather than appending all discovered rules into a monolithic prompt, VAPOR builds a hierarchical expert prompt system over dual visual-semantic clusters, routing each input to specialized guidance matched to its visual regime. On three multimodal benchmarks with ambiguous visual decision boundaries, VAPOR achieves up to 12% absolute gains over established APO baselines. Our analysis shows that visual reflection produces more compact, stable, and interpretable prompts, reframing APO for VLMs as a problem of grounding language search in visual evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.