acceptodds
Under review as a conference paper at ICLR 2027

SAPS-COT: STEP-AWARE VISUAL GROUNDING WITH DYNAMIC RESIDUAL INJECTION FOR INTERPRETABLE MULTIMODAL REASONING

Abstract

While Large Vision-Language Models (LVLMs) enhanced by Chain-of-Thought (CoT) prompting have significantly advanced multimodal reasoning, they predominantly suffer from a “visual black box” dilemma and a “perceptual lock-in effect.” Existing paradigms typically extract visual features once and freeze them, forcing the model to rely heavily on language priors during intermediate reasoning steps. This often leads to visual hallucinations and cascading errors. To address this, we propose SAPS-CoT, a two-stage reasoning framework that seamlessly integrates Step-Aware Patch Selection (SAPS) with a Dynamic Residual Injection mechanism. In the discovery phase, SAPS establishes a fine-grained, explicit alignment between textual reasoning steps and visual regions by leveraging multi-dimensional metrics, including semantic consistency, gradient saliency, and location bias. In the refinement phase, we introduce an in-place latent residual injection mechanism as an alternative to token concatenation. It dynamically updates the model’s visual memory while preserving the original visual-token sequence length. Controlled comparisons show modest but consistent accuracy gains over concatenation, with a small implementation-dependent latency premium. Extensive experiments on complex real-world scenarios demonstrate that SAPS-CoT mitigates visual hallucinations and outperforms state-of-the-art multimodal CoT methods in both accuracy and interpretability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.