U2: Can Models Learn to Ask the Unseen and the Unsolvable?
Abstract
As models become increasingly capable, recent benchmarks have introduced challenging questions that require models to perform complex reasoning and utilize external tools, such as web search or coding sandboxes. However, these benchmarks are costly to construct and typically demand significant effort from human experts for question design and annotation. Consequently, these datasets are difficult to scale up and update over time, posing a challenge as models risk training on similar, static questions. To address this problem, we propose training a question generator that takes academic materials, such as research papers, as input and outputs complex questions requiring a comprehensive understanding of the source text. Specifically, we employ reinforcement learning (RL) and design several rewards to incentivize the model to generate increasingly difficult questions. The most critical of these is the “difficulty reward,” which utilizes another powerful model to answer the generated questions during rollout. This solver model yields an answer distribution, where higher uncertainty indicates a more challenging question. Following training, we employ the generator to construct the Unseen and nsolvable mark (), composed of three 500-question tracks: (1) open-ended questions requiring comprehensive, citation-grounded reports; (2) fact-heavy non-STEM QA requiring extensive searching and reasoning; and (3) calculation-heavy STEM QA requiring source-grounded derivation and computation. Across these fixed evaluation sets, frontier models equipped with external tools retain substantial room for improvement. We further introduce , a fully self-evolving setting in which the question generator, difficulty model, and target solver are all derived from the same base model. The generator produces questions that the base model itself finds difficult, and the validated questions are then used to train a fresh solver initialized from the same base model. The resulting GPT-OSS-20B model improves over the base by 4.8, 3.3, and 3.6 points on GAIA, FRAMES, and HLE.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.