acceptodds
Under review as a conference paper at ICLR 2027

Fix the Reward or Fix the Data? CoSEC, a Counterfactual Self-Evolving Curriculum for Grounded Vision–Language Reinforcement Learning

Abstract

Reinforcement learning with verifiable rewards is widely used to post-train vision–language models (VLMs). Because it scores only whether the final answer is correct, it does not separate answers that rely on the image from answers that rely on prompt templates, option priors, or rendering artifacts. Counterfactual grounding rewards address this by perturbing the visual evidence a question depends on and measuring whether the model's confidence in its answer moves. In current practice the signal is used once, to shape a reward on a fixed training set; the per-sample diagnostic does not feed back into the next round's data, and a shaped reward pinned to one batch leaves the policy room to exploit it. We introduce CoSEC, a Counterfactual Self-Evolving Curriculum in which a single set of counterfactual interventions drives both loops of grounded VLM RL: a self-consistent grounding reward, and a closed-loop curriculum that regenerates constructively-labeled data every round with no human annotation and no real training images. Our central finding concerns the interaction between the two. A grounding reward's effect reverses sign with the data regime, acting as a hazard on fixed data and as a net gain once the data is renewed, so the curriculum is what makes a shaped reward safe to optimize. Containing this failure also improves accuracy: across five VLM scales in two families, CoSEC raises out-of-distribution macro accuracy over outcome-only RL, and the advantage grows when output length is held fixed, indicating that the gain reflects grounding rather than verbosity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.