acceptodds
Under review as a conference paper at ICLR 2027

Counting Rounds Is Not Measuring Progress in Day-Scale Agent Execution

Abstract

Longer agent runs produce more interactions, but what do those interactions add? We argue that neither work-order counts nor instruction labels can stand in for realized progress, and support this claim with three linked analyses of 93 day-scale coding episodes. First, 391 runtime attempts reduce to 283 distinct work orders; GPU-unavailable tasks contribute 83.7% of later orders but only 31.1% of recorded input tokens, exposing a strong mismatch between order counts and resource exposure. Second, two isolated runs of the same judge model agree that only 24 of 61 full-text extension instructions require behavioral change. Third, removing one selected recorded patch from each of six games reproduces behavioral failures that the retained patches prevent under matched conditions. Four of these repairs arose under extensions without a mandatory behavioral change, and one under a fallback continuation. The experiments separate repeatable patch-local effects from instruction semantics: a request to verify can lead to a repair, while another order need not mark another increment. We derive collection and evaluation implications from this distinction and preserve the underlying boundaries in reusable action, handoff, and next-order records.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.