When Does Data Leakage Translate into Evaluation Gains? A Study of VLM Post-Training
Abstract
Why does evaluation-data exposure yield substantial score gains in some post-training settings but only limited average changes in others? We study train–evaluation overlap in vision–language models (VLMs) across training objectives and task settings. On natural tasks, increasing overlap from 0% to 100% raises fixed-set accuracy by 21.2 percentage points for supervised fine-tuning and by 14.0 and 14.2 points for DPO and PPO, but by only about 3.7 points for both GRPO and RLOO. In the adaptive visual-token setting, the five post-training objectives instead show fixed-set changes from -0.7 to +1.3 percentage points, including a near-zero response for supervised fine-tuning. Limited average responses therefore occur in distinct post-training settings and cannot be assigned to an optimizer category alone. To analyze this variation, we combine local propagation analysis with an exact finite-horizon decomposition connecting exposure-induced training differences, their propagation through later updates, and whether the resulting behavioral changes affect the evaluation score. This framework distinguishes a change in training from an observable evaluation benefit without assuming that similar terminal responses share one mechanism. Together, the results show that exposure responses must be interpreted in relation to the training setting instead of being inferred from an optimization label when considered in isolation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.