Fix the Reward or Fix the Data? CoSEC, a Counterfactual Self-Evolving Curriculum for Grounded Vision–Language Reinforcement Learning
Abstract
Reinforcement learning with verifiable rewards is widely used to post-train vision–language models (VLMs). Because it scores only whether the final answer is correct, it does not separate answers that rely on the image from answers that rely on prompt templates, option priors, or rendering artifacts. Counterfactual grounding rewards address this by perturbing the visual evidence a question depends on and measuring whether the model's confidence in its answer moves. In current practice the signal is used once, to shape a reward on a fixed training set; the per-sample diagnostic does not feed back into the next round's data, and a shaped reward pinned to one batch leaves the policy room to exploit it. We introduce CoSEC, a Counterfactual Self-Evolving Curriculum in which a single set of counterfactual interventions drives both loops of grounded VLM RL: a self-consistent grounding reward, and a closed-loop curriculum that regenerates constructively-labeled data every round with no human annotation and no real training images. Our central finding concerns the interaction between the two. A grounding reward's effect reverses sign with the data regime, acting as a hazard on fixed data and as a net gain once the data is renewed, so the curriculum is what makes a shaped reward safe to optimize. Containing this failure also improves accuracy: across five VLM scales in two families, CoSEC raises out-of-distribution macro accuracy over outcome-only RL, and the advantage grows when output length is held fixed, indicating that the gain reflects grounding rather than verbosity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.