Scentinel: Smells Like Trouble for Vision-Language Models
Abstract
Vision-language models (VLMs) often receive auxiliary text describing conditions a camera cannot capture. Whether their use of such text reflects cue semantics, generic prompt sensitivity, or calibrated evidence weighting remains unclear. We introduce SCENTINEL, a benchmark that holds an image and question fixed while varying a textual odor description across five conditions: no odor, congruent, incongruent, irrelevant, and adversarial. Across 1,200 visually under-specified items in eight domains and seven VLMs, the original evaluation pipeline reports pooled irrelevant-controlled judgment shifts of 27.4 percentage points for incongruent cues and 30.0 points for adversarial cues. Explicit instructions to ignore odor reduce these contrasts by 89.8% and 83.7%, respectively, while delayed cues retain most of the standard-profile contrasts. Models often explicitly cite odor in their reasoning. These findings characterize judgment sensitivity across the evaluated cue pools and its substantial suppression by explicit instruction. The benchmark combines paired behavioral contrasts with an audit of stated evidence sources. Supplementary analyses document directional summaries, cue composition, classification provenance, and response validity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.