Accepted but Unfinished: The Salvage Value of Truncated Rollouts in RLVR
Abstract
A policy trained with verifiable rewards can learn to place an answer its verifier accepts and then keep generating until the response cap. The reason is the default scoring path: the four open RLVR stacks we inspected grade whatever text is retained at the cap, so a truncated rollout is not scored zero but earns its salvage value, the mean reward paid to rollouts that hit the cap. This one unreported quantity governs what the policy learns, what the pass rate measures, and what the update does. Learning: in one default-grader run all 848 easy-stratum rollouts are accepted while 11 complete; in another, completed-and-correct output falls from 0.70 to 0.015 in a single update while two thirds of truncated outputs still earn full credit. Measurement: salvage value is exactly recoverable from three logged scalars, verified per difficulty stratum across five policy families; since it enters the pass rate as , the same policy is ranked differently at different caps without training (44.7 pp swing, rank correlation 0.60). Update: salvage is a distinct credit term, the covariance of reward with policy score within truncated rollouts; a preregistered matched fork that pins truncated reward at the checkpoint's own salvage value finds the default-grader arm truncating more in four of six seeds. The published treatments of overlong outputs act on different terms of the same decomposition: a fixed high reward costs completion in all five seeds, masking costs accuracy in nine of twelve matched cells, and neither substitutes for the other. Salvage value is the coordinate along which acceptance and completion separate; it is readable from existing logs and belongs next to the pass rate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.