acceptodds
Under review as a conference paper at ICLR 2027

ReTreVal: Reasoning Tree with Validation and Cross-Problem Memory for Large Language Models

Abstract

Every existing inference-time reasoning framework discards failure context at problem boundaries, leaving a model solving problem 500 no wiser than it was on problem 1. We present ReTreVal (Reasoning Tree with Validation), a training-free framework combining adaptive tree exploration, tool-augmented refinement, typed-failure backtracking, dual-critique scoring, and persistent self-rewriting memory. On the full MATH-500 benchmark, ReTreVal achieves 85.8% pass@1 with ground-truth-assisted verification and 84.2% (421/500) without reference answers in the inference loop. Under label-free verification on MMLU-Pro, it reaches 54.4% (544/1000), compared with 39.1% for Self-Refine. On full-scale label-free MATH-500, cross-problem Reflexion and Self-Refine with persistent memory reach 81.0% and 83.2%, respectively. Across three full-scale reshuffles, MATH-500 varies from 84.8% to 86.2% (1.4 points) and MMLU-Pro from 53.6% to 54.9% (1.3 points), providing no evidence of a large sequence effect in the tested runs. These results position ReTreVal as an auditable structured-inference procedure that performs targeted recovery and retains bounded failure experience without modifying model weights.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.