Says Who? We Read and We Write the Decisions of Vision-Language Models
Abstract
Vision-language models (VLMs) combine visual evidence with instructions, yet it is hard to see how they reach a decision or which parts of the picture drive it. We introduce SimonSays, a benchmark of controlled visual counterfactuals for studying authority attribution, evaluation awareness and sandbagging, and V-Lens, a pair of cheap tools that read and attribute a VLM's decision at the answer position. V-Read applies a closed-form linear shortcut from each layer to the final state, and V-Write applies integrated gradients to image-token states from the picture's own mean state. Our contribution is not the maps themselves but placing them at the answer position of a VLM, testing them against matched image edits, and using them on safety-relevant behaviour. V-Read fits in 7 to 20 minutes, against 2.6 to 8.9 hours for the Jacobian lens, reads the answer earlier than the logit lens on all four models we test, and is the best reader on authority tasks more often than any other. V-Write localises COCO objects as well as KernelSHAP when both are scored on the same regions, at about 1% of its runtime, and its top-ranked image tokens remove more of the answer than those of any compared method on all four models. With SimonSays, V-Lens separates the cues a model registers from those it acts on. Both models register a sign's carrier and an instruction to sandbag without following them, while the smaller model does act on a picture dressed as a test. In a model fine-tuned to sandbag when a small watermark is present, which lowers held-out accuracy from 0.60 to 0.30, V-Write finds the watermark from the answer alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.