acceptodds
Under review as a conference paper at ICLR 2027

Grounding on Request: Learning Modality Control in Vision-Language Models

Abstract

When an image and its caption disagree, can a user reliably control which one a vision-language model (VLM) follows? Accuracy alone cannot answer this question, since a model might give the correct answer from parametric knowledge without following the requested evidence. We therefore evaluate modality grounding control through conflict-preserving counterfactual interventions. Specifically, Flip measures whether an initially correct answer follows an edit to the requested modality, while Stay measures whether it remains stable when the competing modality is edited. To learn this control, we build on neologism learning by introducing new words for visual and textual grounding, training only their embeddings while keeping the pretrained model fixed. We also introduce Counterfactual Twin-Margin (CTM) supervision, which uses paired interventions to teach how answer probabilities should change when the requested evidence changes. Across three VLMs on three existing benchmarks and our multi-attribute Ground-CF benchmark, we find that learned modality words improve control over natural-language instructions. Surprisingly to us, a single learned embedding per modality provides strong lightweight visual control and remains stable on held-out attribute values. LoRA remains strongest overall, but its visual advantage narrows on the held-out evaluations. Moreover, counterfactual supervision improves control in several settings where standard supervision is weaker. Finally, we test our method on real-world news video, and find that visual control improves when the relevant evidence is made easier to access. Overall, our results show that reliable modality control requires testing not only whether an answer is correct, but whether it continues to follow edits to the modality of choice.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.