acceptodds
Under review as a conference paper at ICLR 2027

TRIP-Adv: Verifier-Anchored Token Credit Assignment via Controlled Visual Interventions

Abstract

Reinforcement learning with verifiable rewards improves multimodal reasoning, but most methods assign the same sequence advantage to every response token. This uniform credit treats grounded decisions, language priors, and accidental guesses alike, weakening supervision for tokens that truly depend on visual evidence. We propose TRIP-Adv, a framework that identifies these tokens through controlled visual interventions and assigns credit. An exponential moving average teacher scores each trajectory under the original image, target enhancement, target corruption, and matched random corruption. These comparisons estimate whether target evidence is sufficient to support a token and necessary for its prediction, while reducing sensitivity to generic image changes. Agreement with the outcome verifier and consensus across interventions calibrate the evidence. The calibrated signal is sparsely routed to reshape token advantages and concentrate learning on visually grounded reasoning. Across seven benchmarks, TRIP-Adv achieves the best overall performance among models of comparable scale. The code is available in the appendix.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.