acceptodds
Under review as a conference paper at ICLR 2027

The Umbra Effect: Use–Elicitation Dissociation in Vision-Language Models

Abstract

Anthropic's J-lens study links verbalizable representations to a global workspace in language models. This raises a complementary question for vision-language models: when visual information influences an answer, can an intervention also elicit that information in words? We identify the Umbra effect, a dissociation between causal use and verbal elicitation of visual information, and characterize it using mirrored-image interchange, linear probes, and calibrated signed activation injection. Across five checkpoints from two model families, we characterize three findings: 1) Use–elicitation dissociation. Four Qwen checkpoints show weak spatial pooled-direction elicitation in tested settings where interchange changes answer margins, despite effective color controls. 2) Vision-to-answer handoff. A detailed Qwen2.5-VL-3B study localizes a crossing of vision- and answer-position interchange effects at L23, within a band comprising six tested layers with causally influential, linearly decodable spatial representations. 3) Payload-dependent elicitation. LLaVA-OneVision responds to position-preserving spatial payloads where pooled payloads fail, identifying a boundary on the cross-model pattern. Photograph and clutter tests additionally extend the use–elicitation comparison within the 3B case study. Together, these results motivate auditing causal contribution, decodability, and controlled elicitation as distinct measurements, with payload and dose calibration guiding the interpretation of negative readouts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.