acceptodds
Under review as a conference paper at ICLR 2027

How Vision-Language Models Form Ambiguous Depth Interpretations: A Sparse Feature Analysis of Visual Evidence Transformation

Abstract

Vision-language models (VLMs) can produce different depth interpretations of ambiguous images, yet how conflicting visual cues are progressively organized into interpretation-related representations within visual encoders remains poorly understood. Here, we investigate this question using Mach card stimuli containing convex-supporting texture cues and concave-supporting contour cues. We construct controlled stimuli with isolated cues and consistent depth conditions, and analyze sparse visual features across CLIP layers of the LLaVA visual encoder using sparse autoencoders (SAEs). By combining cross-condition feature selection, layer-wise tracking, spatial localization, and causal intervention, we characterize the relationship between visual cue encoding, depth interpretation representations, and model outputs. We find that features associated with specific visual cues emerge in early visual layers, whereas features representing opposing depth interpretations mainly arise in intermediate layers and can be jointly activated under ambiguous conditions. Feature steering and ablation further demonstrate that manipulating these candidate features systematically alters the model’s convex–concave logit margin, with selected contour related interventions shifting the model’s preferred depth interpretation. These findings suggest that ambiguous interpretation in VLMs does not arise from a direct mapping between visual input and final output, but through a hierarchical transformation from visual cue encoding to depth-related interpretation representations. Our study reveals the internal organization of ambiguous visual interpretation in vision-language models and provides a mechanistic framework for understanding how multimodal systems transform competing visual evidence into language-level decisions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.