acceptodds
Under review as a conference paper at ICLR 2027

Rewind to Rethink: Localized Repair for Multimodal Reasoning

Abstract

Multimodal reasoning models can continue to rely on an early judgment even after generating additional reasoning, allowing an initial error to persist into the final answer. We define step-to-step interventional support using offline Shapley attribution: preceding steps are retained or masked, and their marginal contributions to the mean log-likelihood of fixed subsequent text are averaged across coalitions. The resulting matrix reveals how an early judgment can continue to support later reasoning and the final answer, a pattern we call answer inheritance. Inspired by this observation, we propose BANKROLL, a training-free method that rolls back to regenerate an earlier judgment together with the reasoning built on it. Cashout monitors intermediate answers to determine when thinking stops. Refinance uses online local masking tests as a heuristic to select a rollback point, truncates the chain before that step, and samples alternative continuations. Writeoff retains the original chain if no rollback point is chosen, no continuation is selected, or the probe-score acceptance check fails. Experiments across five benchmarks and three multimodal models show that BANKROLL consistently improves task performance over standard chain-of-thought (CoT).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.