acceptodds
Under review as a conference paper at ICLR 2027

AnchorPlay: Reliable Self-Play for Video Grounding

Abstract

Self-play can expand training supervision for video grounding without additional human annotations, but it also risks reinforcing its own localization errors: agreement among predictions does not establish that an event's boundaries are correct. We introduce \ours, an end-to-end self-play framework that jointly optimizes a shared model as a questioner and a grounding solver. The central idea is to decouple question evolution from the construction of localization targets. A frozen reference model proposes evidence anchors, whose reliability is assessed by comparing predictions across transformed views after mapping them into a common coordinate system. Conditioned on accepted anchors, the questioner learns to generate queries matched to the solver's capabilities, while the solver learns to localize the corresponding objects or events. Both roles are updated in the same optimization step, allowing the curriculum to evolve around targets supplied by a reference model independent of the evolving policy. Reliability-weighted reinforcement learning emphasizes better-supported targets, while grounding-preservation regularization encourages retention of reliable localization behavior from the initial model. Experiments on temporal and spatio-temporal grounding benchmarks show improvements over initial checkpoints across both general-purpose and task-adapted backbones. Although self-play training uses only grounding tasks, \ours also improves performance on additional video and image understanding tasks, demonstrating transfer beyond its training objectives.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.