acceptodds
Under review as a conference paper at ICLR 2027

Beyond Fixed Supervision: Alternating Self-Play for Region-Grounded Visual Retrieval-Augmented Generation

Abstract

Region-grounded visual retrieval-augmented generation models are typically trained on fixed datasets of human-annotated questions, answers, and evidence regions. Such supervision cannot adapt to the model's evolving weaknesses, particularly limiting continued improvement in region-level evidence localization. We present Self-Play Visual RAG, a framework comprising a proposer, a solver, and an environment. The proposer generates questions, reference answers, and evidence regions directly from visual documents; the environment filters malformed or unsupported tasks; and the solver learns from accepted tasks. Rather than selecting tasks according to a predefined ability boundary, our final strategy prioritizes valid tasks that the current solver fails, automatically constructing a hard-task curriculum without additional human annotations of question-answer-evidence-region tuples. We evaluate on a document-disjoint set of 317 manually reviewed questions constructed from SlideVQA and ViDoSeek source pages. The final solver achieves 80.44% Soft EM and 22.74% [email protected], exceeding the best evaluated general-purpose vision-language baseline for each metric by 2.52 and 15.53 percentage points, respectively. The proposer also improves through self-play, reaching an 81.70% joint environment acceptance rate, 11.30 percentage points higher than the strongest zero-shot proposer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.