acceptodds
Under review as a conference paper at ICLR 2027

ClueJigsaw: Learning to Reason with Visual Clues through Jigsaw Reconstruction

Abstract

Visual reasoning is essential for advancing vision-language models (VLMs) toward a deeper understanding of the visual world. However, perceiving and interpreting visual inputs remains challenging when familiar spatial structures are disrupted, as exemplified by jigsaw reconstruction. We find that successful reconstruction benefits from the joint use of global semantics and local structural clues. Motivated by this finding, ClueJigsaw is proposed as a post-training framework for clue-aware jigsaw reconstruction. We introduce a clue-guided trajectory construction method and use SFT cold-start training to develop models’ ability to independently infer semantic clues and structural clues from shuffled inputs. Clue-aware reinforcement learning is then introduced through an easy-to-hard curriculum of reconstruction subtasks to progressively improve clue identification and utilization. We design a two-part reward that combines answer correctness with VLM-based assessment of clue use, additionally rewarding reasoning grounded in appropriate visual clues. Experiments show substantial improvements on both and jigsaw tasks, with reconstruction accuracy rising from 10.12% to 97.38%, alongside gains on general visual reasoning benchmarks. These results highlight the value of visual clue supervision and verification in VLM post-training, helping models connect visual understanding with clue-grounded reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.