acceptodds
Under review as a conference paper at ICLR 2027

CAVE-RL: CONFIDENCE-ADAPTIVE VISUAL EVIDENCE FOR MULTIMODAL REINFORCEMENT LEARNING

Abstract

Reinforcement learning (RL) for multimodal models increasingly relies on learned visual verifiers, but those verifiers are imperfect precisely on the visually ambiguous inputs where policy optimization is most likely to exploit them. We propose CAVE-RL (Confidence-Adaptive Visual Evidence RL), a reward construction that combines sparse exact task feedback with visual proxy rewards whose influence is attenuated by predictive entropy and ensemble disagreement, then anchors policy updates to a frozen reference policy. The method is architecture-agnostic and can be inserted into PPO/GRPO-style post-training whenever an action can be mapped to a visual claim scored by one or more verifiers. We construct 30,555 multimodal image-instruction instances from 1,797 real handwritten-digit images, covering five language-conditioned transformation tasks, held-out prompt paraphrases, visual distribution shifts, 60% corruption in verifier training labels, and five independent random seeds. Only 15% of RL rollouts reveal the exact task outcome. Under this budget, CAVE-RL improves the five-condition robust average by 2.17 pp over SFT and 0.70 pp over sparse-reward PPO, while reducing expected calibration error from 12.49% to 8.34%. It recovers 74% of the robustness gain of an oracle PPO baseline that observes the exact outcome on every rollout.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.