Reliable Local Credit from Stratified Continuation Evidence
Abstract
Sample-based local credit uses auxiliary continuations to evaluate intermediate reasoning decisions. With a finite evidence budget, accurate value estimates are only part of the problem: their errors interact as local credits are converted into a policy update. We characterize this evidence-to-update relationship and show how reliability depends on both evidence covariance and update geometry. We develop Stratified Segment Policy Optimization (S-SPO), which changes the joint design of auxiliary continuations while preserving the operational value target, local-credit rule, and continuation count per query. Under an explicit independence and reuse contract, equal-mass stratification guarantees non-increase of every fixed linear quadratic update risk; a first-order guarantee with a controlled remainder extends the analysis through normalization. At nine continuations per query, S-SPO improves full-parameter frozen-gradient reliability. Across three matched RhoMath-1.1B training seeds, it raises fixed-endpoint validation accuracy by 2.75 percentage points on average and official GSM8K test accuracy by 2.26 points, with every seed pair positive on both evaluations. A second-model comparison on Llama-3.2-1B-Instruct yields a 2.85-point validation gain. These results connect the design of finite continuation evidence to reliable local credit and improved learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.