acceptodds
Under review as a conference paper at ICLR 2027

Reading Between the Hands: Diagnosing Visual Measurement in Vision-Language Models

Abstract

Vision-language models can solve increasingly complex multimodal tasks yet remain unreliable at reading analogue dial instruments. Clocks provide a familiar and highly controlled testbed for studying how fine-grained visual geometry is transformed into numerical readings. Rather than optimising benchmark performance, we ask *when* and *where* models fail at clock reading. Across four VLMs from the Qwen, Molmo and Llama families with distinct multimodal architectures, we combine controlled visual perturbations, equivalent text-based representations of hand geometry, test-time reasoning, component-wise supervised fine-tuning, and layer-wise probes of hour-, minute-, and second-hand geometry. Taken together, these analyses highlight three important aspects of the problem. First, clock reading does not behave like a simple perceive-then-compute pipeline: models can perform substantially better from images than when ground-truth hand positions are supplied explicitly in text, while additional reasoning can recover the text-based mapping but degrade image-based reading. Second, hand geometry is nevertheless linearly recoverable across the visual and language pathways, showing that the availability of geometric information does not guarantee its successful downstream use. Third, there is no universal repair point: the most effective adaptation target differs across architectures, and large improvements in clock reading rarely transfer to eight related dial instruments. Together, these results show that improved task performance, accessible geometric representations, and transferable measurement ability are distinct properties. We therefore characterise clock reading as an architecture-dependent problem of coordinating geometric perception with numerical readout, rather than as a single missing perceptual or reasoning capability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.