TRIP-Adv: Verifier-Anchored Token Credit Assignment via Controlled Visual Interventions
Abstract
Reinforcement learning with verifiable rewards improves multimodal reasoning, but most methods assign the same sequence advantage to every response token. This uniform credit treats grounded decisions, language priors, and accidental guesses alike, weakening supervision for tokens that truly depend on visual evidence. We propose TRIP-Adv, a framework that identifies these tokens through controlled visual interventions and assigns credit. An exponential moving average teacher scores each trajectory under the original image, target enhancement, target corruption, and matched random corruption. These comparisons estimate whether target evidence is sufficient to support a token and necessary for its prediction, while reducing sensitivity to generic image changes. Agreement with the outcome verifier and consensus across interventions calibrate the evidence. The calibrated signal is sparsely routed to reshape token advantages and concentrate learning on visually grounded reasoning. Across seven benchmarks, TRIP-Adv achieves the best overall performance among models of comparable scale. The code is available in the appendix.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.