Counterfactual Visual On-Policy Distillation
Abstract
Visual on-policy distillation uses privileged evidence crops to improve a full-image student's fine-detail recognition. However, distilling teacher outputs on individual images does not explicitly distinguish how the target detail and surrounding content contribute to answer preferences. Our idea is to supplement this supervision with controlled detail changes, observing how teacher preferences respond while holding the question and surrounding scene as constant as possible. We propose CF-OPD, counterfactual visual on-policy distillation, which uses semantic replacement to construct verified image pairs with different correct answers. At each student-generated prefix, CF-OPD compares the two crop-conditioned teacher distributions, using their contrast to adjust a student-anchored distribution before mixing it with the current crop teacher. Independent rollouts on both images let each side learn from this paired supervision. We also introduce CF-Bench-X, an evaluation benchmark comprising 595 verified semantic image pairs. Its Flip rate measures the proportion of pairs answered correctly on both images; its Blind rate measures unchanged predictions despite a verified answer change. CF-OPD achieves the highest average accuracy across six visual benchmarks among the compared Qwen3.5-based methods at both scales. At 9B, it improves recognition and paired evaluation over independent visual OPD on the same inputs. Code and dataset will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.