TraceStencil: Signed Temporal Neighborhoods for Token-Level Policy Optimization in RLVR
Abstract
Reinforcement learning with verifiable rewards supplies sequence-level feedback while updating token-level probabilities, leaving the choice of which neighboring positions should influence each update largely implicit. We introduce TraceStencil, a token-level policy-optimization objective that represents this choice as a signed temporal neighborhood. Over the selected past and/or future offsets, TraceStencil aggregates likelihood ratios in separately clipped directional branches and applies their detached product to the current-token score, making direction, radius, boundary exposure, and gradient flow explicit. We derive exact structural relations among these quantities and identify the conditions under which the objective reduces to a forward-only trace estimator. Controlled multi-run experiments on mathematical reasoning compare trace-free, one-sided, and balanced geometries under matched training and evaluation. In the pooled six-run comparison, the balanced geometry is higher on the primary coverage@256 endpoint and secondary pass@1 than its future-only counterpart, with reported pooled intervals excluding zero. Across additional model scales and mathematical benchmarks, balanced geometry yields the higher point estimate in every evaluated setting, and the larger Qwen3 configuration has a positive interval. TraceStencil therefore provides both a practical optimization rule and an auditable framework for studying how sequence-level evidence is allocated across autoregressive token updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.