acceptodds
Under review as a conference paper at ICLR 2027

Seeds Do Not Replicate Runs: The Unit of Replication, the Estimand and the Budget Rule in Agent RL

Abstract

Agent RL reports a credit-assignment comparison as “A beats B by over three seeds”. Three choices hide in that sentence, and on one pool of runs we measure what each decides. First, a seed argument is not a unit of replication. Whether re-running it reproduces the run is decided below the library: of the TRL and verl-agent configurations we census, only in-process generation replays to the bit, on one model and task, and our own loop replays only with its nondeterministic kernels off. Within one launch, our registered designs leave the seed's share of run-to-run variance unresolved, with upper ends near a quarter. Across launches, a re-executed seed argument moves the endpoint as far as changing it. On the released verl-agent stack, sixteen seed arguments put the seed's share at , at most , at the final evaluation. Second, the size of a horizon effect is a property of the estimand and the budget rule. On KnobChain, a corridor whose length we set, 384 paired runs of Qwen2.5-7B-Instruct with a turn-level credit rule against an episode-level one show the endpoint gap shrinking nearly threefold with the horizon under a budget that grows with it. The shrinkage belongs to the reading: the method arm sits at the ceiling of the success scale in five of six depths; at a fixed update count the gap does not shrink; the rate at which runs first reach the success threshold shows no resolvable trend; and off the ceiling, at constant anchor coverage, the endpoint gap grows, factor . On the field's own stack the endpoint and hazard readings differ in size, log-difference . That experiment reproduces the published growth under both estimands. The separation holds at the registered threshold and not below it. Third, three runs per arm resolve neither the effect nor the horizon trend, and the effect itself is not in question: under this dense shaped reward the method multiplies the takeoff hazard by two to three at every horizon. Report replications in runs rather than seed arguments, name the estimand and the budget rule, and fix the success threshold before training. The environment and the code accompany this submission; the run pool and the registration record will be released with the paper.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.