Measuring and Amplifying Self-Verification on Hard Problems
Abstract
An open question in scalable oversight is whether language models can assess solutions to problems that they cannot reliably solve. We test this by separating an assessor's informedness from its leniency. Informedness refers to a binary assessor's ability to distinguish correct from incorrect solutions. Leniency refers to its tendency to accept. Across six language models and seven benchmarks in mathematics, coding, and knowledge, we evaluate assessment on problems the model itself rarely solves (own solve rate ). We find that informedness on difficult problems strongly depends on task family: it is significantly higher than chance on mathematics and coding, but near the chance baseline on knowledge tasks. Next, we investigate committees of assessors and find that they can increase informedness compared to a single assessor. However, the gain is not automatic: with majority voting, the committee is worse than a single assessor in 6 of 18 settings. We therefore apply a simple mathematical model to predict committee behavior and select consensus thresholds. Taken together, our results identify task family, assessor behavior, and aggregation rule as key determinants of whether oversight remains effective beyond a model’s own solving capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.