WHEN MLLMS MISS WHAT HUMANS READ: DIAGNOSING AND LEARNING TO READ SQUINT TEXT
Abstract
Multimodal LLMs can fail to read text embedded in an image’s coarse structure— text a human recovers by squinting—and report that no text is present. We name this class Squint Text and study it as a recognition deficiency, motivated by cases from a production moderation pipeline. We introduce SQUINTBENCH, compris- ing 1076 production-derived moderation images, and SQUINTCTRL, a controlled synthetic suite for mechanism diagnosis. Frozen-weight analysis shows that hid- den text leaves a coarse signal while glyph identity is weakly represented; iden- tity becomes substantially more accessible when the correct spatial regions are supplied, implicating spatial selection as a major bottleneck. Suppressing high- frequency content also improves reading at a fixed visual-token budget, while vi- sual token count may also contribute. These findings motivate SQUINT OPSD: a frozen teacher receives a downscaled view and the ground-truth transcription as privileged context, and its token-level distributions supervise a student along on-policy trajectories. At inference, the student reads the original image alone. SQUINT OPSD reaches 88.61 ANLS on SQUINTBENCH, versus 26.50 for the unadapted backbone, 24.44 for the best general-purpose model on the original view, and 68.60 at best for an off-the-shelf model under a downscaled view. These results substantially narrow the gap between available visual evidence and its gen- erated read-out without auxiliary views at inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.