acceptodds
Under review as a conference paper at ICLR 2027

Two Deaths of Binary Reward: Arrival and Escape in Formal Language Tasks

Abstract

Initial correctness does not determine how effectively binary-reward reinforcement learning adapts after obtaining a correct sample. We separate first arrival from subsequent escape using formal language tasks with exact verification and controlled output factorization. Factored policies with identical initial correctness, entropy, reward variance, and output-space cardinality exhibit a 49.5 spread in median post-arrival escape time near the arrival boundary. An exact first-update calculation for no-baseline REINFORCE with plain SGD identifies dependence on factorization and loss normalization. Under isolated updates that leave the policy unchanged before the first correct sample, learning-rate changes preserve the arrival distribution. On 30 regex instances drawn uniformly from 132 buildable members of a 200-instance corpus draw, tripling the learning rate of a pretrained 0.5B model increases final success from 9/90 to 27/90 runs while preserving first-arrival episodes in all 41 arriving pairs. Shared-prompt updates allow learning-rate changes to affect arrival as well. In a 30-prompt setting with three shared training trajectories per arm, two rates both yield 0/39 final successes, while mean final accuracy differs from 0.010 to 0.156; a higher rate or longer continuation recovers success on a subset. At the tested probe depth, zero observed hits do not certify rare arrival during training. These results distinguish obtaining correct samples from converting them into successful learning, and identify sufficient update conditions under which intervention effects on the two stages separate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.