Recursive Self-Verification: Progressive Binary Verification for Language Model Self-Evolution
Abstract
Self-evolving language models learn from their own generations by using the model to identify responses that yield useful training data. However, directly assigning fine-grained, multi-level scores to responses is often unreliable, particularly for smaller models that struggle with subtle quality distinctions, resulting in unreliable preference pairs. To address this limitation, we investigate whether fine-grained self-verification can be achieved by decomposing complex evaluations into a sequence of simpler binary decisions. Specifically, we introduce Recursive Self-Verification (RSV), a concise yet effective framework that evaluates responses through sequential binary verification steps ordered from strict to relaxed quality criteria. At each step, RSV uses normalized token probabilities for Yes and No to recursively allocate probability mass across quality levels, deriving the final evaluation score as the expected value of this distribution. These refined scores are then used to rank responses and construct high-quality preference pairs for model optimization. Experiments on the Knights-and-Knaves, TruthQuest, and MATH-Hard benchmarks demonstrate that RSV-guided preference learning consistently outperforms baseline methods in reasoning tasks across various model scales. Notably, on K&K with Qwen3-1.7B, the best RSV setting achieves an average accuracy of 63.1%, outperforming the strongest non-RSV baseline by 8.3 percentage points. These gains also extend to GRPO, confirming RSV's effectiveness under on-policy optimization for language model self-evolution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.