RealNLR: Do VLMs Learn Nonlocal Visual Algorithms from Synthetic Fine-Tuning?
Abstract
Vision-language models (VLMs) do well on standard benchmarks but fail simple tasks that require chaining visual evidence across an image, called nonlocal visual reasoning (NLR): re-identifying a transformed object, following a chain of cues (saccadic search), and tracing a wire through clutter (contour tracing). We ask whether fine-tuning can teach these visual algorithms, and what the model learns if it does. We fine-tune five open VLMs (4B to 32B) with QLoRA on synthetic versions of the three tasks. We vary the adapted side (language or vision), the supervision (answer only, or intermediate steps too), the amount of data, and the rendering style (default or randomized). To measure transfer, we introduce RealNLR, a benchmark that poses the same tasks on realistic images. We find that (i) fine-tuning raises in-distribution accuracy by 13 points on average, and the best runs go from near zero to 98%, but transfer to RealNLR is zero; (ii) process supervision teaches the steps, with InternVL3-14B writing 96% of out-of-distribution search trajectories exactly, yet its answer still stops at the chain lengths seen in training; (iii) internal analyses show that fine-tuning improves one stage of the procedure on synthetic images, but not how the model builds visual state from a natural image; and (iv) randomized-style training does not improve transfer (mean change −0.6 points against a matched control). The models learn the dataset, not the algorithm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.