Beyond Final Outcomes: Process-Aware Relative Distribution Matching for LLM Reasoning
Abstract
Distribution matching has emerged as a promising paradigm for LLM reasoning by shaping probability distributions over complete trajectories to preserve diverse rewarding paths and thus improve reasoning performance. However, existing approaches remain largely trajectory-level, leaving two questions unresolved when incorporating intermediate reasoning information: where along a long reasoning trajectory should process information be measured, and how should it influence the distribution-matching objective? We first examine the former through counterfactual continuations. Building on high-entropy reasoning points as candidate decision boundaries, we find that alternative continuations from these forks exhibit greater variation in downstream verified-answer progress than those from neighboring non-fork positions, supporting their use as process-sensitive boundaries. We further observe substantial heterogeneity across segments: improving and deteriorating transitions coexist within both correct and incorrect trajectories, while a small subset of segments accounts for a disproportionate fraction of progress variation. These observations motivate distinct uses of intermediate progress in distribution matching. Progress magnitude controls credit allocation, reallocating optimization emphasis across segments while preserving the distribution-matching residual, whereas progress direction shapes the target distribution by adaptively tempering the reference-policy prior. Based on this design, we propose **P**rocess-**A**ware Relative **D**istribution **M**atching (PADM). Across trajectories sampled for the same prompt, a group-relative objective aligns relative policy probabilities with the resulting process-aware target. Theoretically, we show that exact relative matching recovers the intended target distribution, while magnitude-based segment weighting preserves this matching condition. Empirically, we evaluate PADM on six challenging mathematical reasoning benchmarks across two model scales, demonstrating consistent improvements over competitive baselines, with gains of up to 4.63 points in average accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.