acceptodds
Under review as a conference paper at ICLR 2027

Self-Evolving Visual Questioner

Abstract

Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, and visually grounded questions remains underexplored. Existing visual questioners' performance is bottlenecked by the availability of high-quality training data and the cost of curating it. We show that a VLM can iteratively improve itself as a visual questioner without external teachers or new human annotations. We propose a self-evolving framework that uses a VLM itself as a proposer, rewriter, and filter to produce harder, more informative, and visually grounded questions, while maintaining diversity across visual intents. These questions are then used to train the VLM in both questioner and answerer modes, with the updated questioner proposing questions for the next round. To evaluate the questioner, we introduce an evaluation protocol that assesses questions along perception, reasoning, and diversity dimensions. Experiments across three backbone VLMs show that our method substantially enhances question quality, increasing the mean question-generation score by approximately 78% after two rounds. Under the same training-data budget, our self-supervision is more effective than training on static source annotations. Moreover, the self-evolving questioner preserves its answering ability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.