When Good Verifiers Go Bad: Silent Negative Transfer in Verifier-Guided VLM Training
Abstract
Verifier reliability is not portable across tasks. A verifier-guided self-DPO pipeline that produces genuine held-out gains on MathVista +9.6 points on the self-training data and +8.0 points held out) can become harmful on MMMU. Crucially, the failure is not visible from the target-task self-training signal: averaged across six learner-verifier configurations, performance on the MMMU self-training data still improves by +3.52 points, while held-out performance drops by 1.42 points. We term this silent negative transfer:a verifier previously validated as useful can continue to exhibit the signatures of successful self-training after its induced update ceases to transfer to unseen data. This raises a more basic question: how can we tell whether verifier supervision is reliable on the target task? A downstream failure alone cannot answer this question, because poor generalization may arise from two distinct sources: the verifier may induce an incorrect learning direction, or the correct learning direction itself may fail to generalize beyond the self-training distribution. We formalize these effects through gradient fidelity F, which measures alignment between the verifier-induced and correct training directions, and gradient transferability T, which measures alignment between the correct training and held-out directions. This decomposition yields a conservative Safe-Transfer Margin: positive first-order held-out alignment is certified when Motivated by this diagnosis, we introduce a simple yet effective precision-first filtering method, Asymmetric Acceptance Gating (AAG), which selects preference pairs according to the verifier's absolute confidence in the response receiving the positive update. Gradient-alignment analyses on MMMU show that AAG increases gradient fidelity from 0.29 to 0.42. In a separate same-cell analysis, the correct training direction remains positively aligned with the held-out direction T=0.31, yet raw verification rotates the induced update to negative held-out alignment -0.13; AAG restores positive alignment +0.22. Across all six MMMU learner–verifier configurations, AAG outperforms raw verifier-guided training on held-out evaluation. These results redefine what it means for a verifier to be reliable: not merely that it has worked before, but that it induces the right update on the task at hand—and that this update remains useful beyond the data that generated it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.