acceptodds
Under review as a conference paper at ICLR 2027

Read Less, Recall More in Document VLMs

Abstract

Document vision-language models (VLMs) fine-tuned on forms see sensitive fields both as pixels and as supervised text, yet whether a field's legibility decides what they memorise is untested: resolution studies change the whole page and its visual-token budget. We introduce Visual Canaries, synthetic identifier–secret documents with matched no-canary controls, and run two interventions on 3 released VLMs: a rendering ladder that changes only the secret's glyphs at a fixed target, exposure and token count, and a fixed-weight test that changes only the inference context. On Qwen2-VL-2B-Instruct, lowering the control's transcription accuracy from 1.00 to 0.11 raises the paired memorisation gap from +0.24 to +3.63 [+3.00, +4.26] bits on a 96-distractor rank scale (+13.25 bits on a fitted tail); a corrupted field with the largest glyphs tracks legibility rather than glyph size, the ordering holds on all 3 backbones, and duplication amplifies it by +3.69 bits per doubling. With weights fixed, restoring the page paired with a secret during training recovers +3.18 [+1.90, +4.46] bits although that page holds none of the secret's pixels; a caller with the identifier, the paired page, the caption format and assistant-prefill capability then samples the secret for 96.0% of Qwen and 90.0% of SmolVLM canaries at one query, and for none without the page. Memorisation risk in document fine-tuning thus depends on what the model can read and on the context an auditor supplies: audits must probe with training-like pages, and redaction must cover both pixels and targets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.