Two Deaths of Binary Reward: Arrival and Escape in Formal Language Tasks
Abstract
Initial correctness does not determine how effectively binary-reward reinforcement learning adapts after obtaining a correct sample. We separate first arrival from subsequent escape using formal language tasks with exact verification and controlled output factorization. Factored policies with identical initial correctness, entropy, reward variance, and output-space cardinality exhibit a 49.5 spread in median post-arrival escape time near the arrival boundary. An exact first-update calculation for no-baseline REINFORCE with plain SGD identifies dependence on factorization and loss normalization. Under isolated updates that leave the policy unchanged before the first correct sample, learning-rate changes preserve the arrival distribution. On 30 regex instances drawn uniformly from 132 buildable members of a 200-instance corpus draw, tripling the learning rate of a pretrained 0.5B model increases final success from 9/90 to 27/90 runs while preserving first-arrival episodes in all 41 arriving pairs. Shared-prompt updates allow learning-rate changes to affect arrival as well. In a 30-prompt setting with three shared training trajectories per arm, two rates both yield 0/39 final successes, while mean final accuracy differs from 0.010 to 0.156; a higher rate or longer continuation recovers success on a subset. At the tested probe depth, zero observed hits do not certify rare arrival during training. These results distinguish obtaining correct samples from converting them into successful learning, and identify sufficient update conditions under which intervention effects on the two stages separate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.