LLM-as-an-Improver: Turning Verification into Better Candidates
Abstract
Verifier-based selection chooses the most promising solution from a pool of candidates generated at inference time. Because it only ranks a fixed candidate pool, however, it must fail when every initial candidate is incorrect, even if verification identifies useful failure evidence. We introduce LLM-as-an-Improver and propose Verify–Repair–Reselect (VRR), which reuses that evidence to expand the candidate set before final selection. VRR retains the initial winner while conditionally generating complementary alternatives, screens invalid and duplicate outputs, and then reselects from the retained winner and the surviving alternatives. Our experiments show that VRR can recover correct solutions even when all candidates in the initial pool are incorrect. These results demonstrate the feasibility of using verification feedback not only as a ranking signal but also as a generation signal for constructing correct candidates beyond the initial pool.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.