acceptodds
Under review as a conference paper at ICLR 2027

TraceStencil: Signed Temporal Neighborhoods for Token-Level Policy Optimization in RLVR

Abstract

Reinforcement learning with verifiable rewards supplies sequence-level feedback while updating token-level probabilities, leaving the choice of which neighboring positions should influence each update largely implicit. We introduce TraceStencil, a token-level policy-optimization objective that represents this choice as a signed temporal neighborhood. Over the selected past and/or future offsets, TraceStencil aggregates likelihood ratios in separately clipped directional branches and applies their detached product to the current-token score, making direction, radius, boundary exposure, and gradient flow explicit. We derive exact structural relations among these quantities and identify the conditions under which the objective reduces to a forward-only trace estimator. Controlled multi-run experiments on mathematical reasoning compare trace-free, one-sided, and balanced geometries under matched training and evaluation. In the pooled six-run comparison, the balanced geometry is higher on the primary coverage@256 endpoint and secondary pass@1 than its future-only counterpart, with reported pooled intervals excluding zero. Across additional model scales and mathematical benchmarks, balanced geometry yields the higher point estimate in every evaluated setting, and the larger Qwen3 configuration has a positive interval. TraceStencil therefore provides both a practical optimization rule and an auditable framework for studying how sequence-level evidence is allocated across autoregressive token updates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.