ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). Yet, reward-driven reasoning can produce unnecessarily long rollouts, repeated verification, and accumulated errors that reduce coherence and exhaust the generation budget. Existing long-context organization methods commonly rely on external compressors, cache mechanisms, or auxiliary modules. We propose ReSum, an RLVR framework that internalizes self-summarization and related reasoning-control behavior to organize subsequent reasoning within the policy. Our pilot studies reveal that summary-like cues occur around high-uncertainty states, are followed by lower token-level entropy, and can improve recovery from incorrect rollout prefixes. Building on these observations, ReSum constructs adaptive contrastive rollout trees around two complementary events: Natural Points (NPs), where a model-generated cue is masked, and Artifact Points (APs), where a cue is injected at a selected non-cue position. ReSum then applies a dual-reference, cue-aware advantage that evaluates each continuation against cue-containing and cue-free rollout distributions, providing state-local credit assignment for beneficial reasoning-control behavior. Across six mathematical benchmarks, multiple backbone sizes and families, and a multimodal geometry benchmark, ReSum consistently improves reasoning accuracy by an average of 4% while reducing rollout length by 18.6%. These results demonstrate a practical RLVR approach for achieving a stronger accuracy–length trade-off in verifiable long-horizon reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.