acceptodds
Under review as a conference paper at ICLR 2027

Bifocal Self-Rewarding: Aligning Answer Sufficiency with Visual Fidelity in VLM Reasoning

Abstract

Outcome-reward reinforcement learning (RL) can improve multimodal reasoning without establishing whether the reasoning used to produce an answer is both sufficient for that answer and faithful to the image. Existing methods address complementary aspects of this problem: some verify the connection between intermediate reasoning and the final answer, while others strengthen the extraction or use of visual evidence. Yet neither direction alone establishes whether the same reasoning chain satisfies both properties. We introduce Bifocal Self-Rewarding, a dual-view framework that evaluates the same reasoning chain along these two dimensions. In the image-stripped view, the current policy uses its own reasoning chain to recover the answer without access to the image, testing whether the chain itself contains sufficient information. In the image-conditioned view, the frozen RL initialization provides a fixed perceptual reference, favoring correct chains that contain concrete visual evidence supported by the image. Using the current policy for sufficiency allows the evaluator to follow changes in the policy's reasoning, whereas freezing the fidelity evaluator provides a fixed reference throughout training. Both rewards are therefore assigned directly to the reasoning chain used to produce the answer, without requiring annotated visual claims or a separately trained judge. Across diverse benchmarks spanning reasoning, perception, extraction, and hallucination, Bifocal consistently improves multimodal reasoning across Qwen2.5-VL and Qwen3-VL backbones from 3B to 8B parameters while preserving visual capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.