Counting Rounds Is Not Measuring Progress in Day-Scale Agent Execution
Abstract
Longer agent runs produce more interactions, but what do those interactions add? We argue that neither work-order counts nor instruction labels can stand in for realized progress, and support this claim with three linked analyses of 93 day-scale coding episodes. First, 391 runtime attempts reduce to 283 distinct work orders; GPU-unavailable tasks contribute 83.7% of later orders but only 31.1% of recorded input tokens, exposing a strong mismatch between order counts and resource exposure. Second, two isolated runs of the same judge model agree that only 24 of 61 full-text extension instructions require behavioral change. Third, removing one selected recorded patch from each of six games reproduces behavioral failures that the retained patches prevent under matched conditions. Four of these repairs arose under extensions without a mandatory behavioral change, and one under a fallback continuation. The experiments separate repeatable patch-local effects from instruction semantics: a request to verify can lead to a repair, while another order need not mark another increment. We derive collection and evaluation implications from this distinction and preserve the underlying boundaries in reusable action, handoff, and next-order records.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.