acceptodds
Under review as a conference paper at ICLR 2027

When Agents Change Their Own Test: Prospective Target Preservation under Candidate-Induced Task Expansion

Abstract

Open-ended and self-improving (autotelic) agents increasingly help construct the tasks, tests, or subproblems on which progress is later judged. This creates a same-candidate dilemma: if candidate update expands the evaluation population, may those candidate-caused additions change the objective used to decide whether itself is accepted? We formalize candidate-induced target endogeneity and distinguish it from adaptive sampling, missing evaluation coverage (support failure), and intentionally performative objectives. Under prospective semantics, the current target distribution is fixed using only pre-candidate information; evidence may still be sampled adaptively and reweighted to that target, while candidate-caused tasks affect only later targets. We prove that any nontrivial same-round target change can reverse some bounded threshold decision even under exact, noiseless evaluation, so greater statistical precision cannot repair a change in the quantity being judged. Exact replay on 384 recorded LLM update-effect vectors yields false-promotion rates (acceptance under the post-candidate target despite rejection under the prospective target) of 73.4% and 54.1% in the primary and independently seeded confirmation phases when only the target semantics change. More importantly, a pre-specified capability-conditioned task-admission analysis on naturally trained Prioritized Level Replay policies in the Procgen benchmark finds 13/48 false promotions across 11/16 StarPilot seeds and, in an independently seeded replication, 36/192 across 29/64 Heist seeds. Eleven Heist reversals have prospective and post-target signs jointly separated from zero under our 95% certificate. Mean disagreement under same-size candidate-independent expansion is 3.2% and 2.5%, respectively. The Procgen CoinRun environment has no primary false promotions, providing a negative boundary. These results identify target timing as a distinct evaluation contract for self-expanding agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.