WHAT HANDOFF LOGS HIDE: PAIRED ACCOUNTING FOR LLM-AGENT WORK
Abstract
Higher success rates. Fewer handoffs. More work left behind. Task success is the dominant metric for tool-using language agents, yet it can misstate the value of autonomy in systems where failed or incomplete work is handed to a stronger agent or human. An additional action can improve the acting agent’s chance of completion while simultaneously making an eventual handoff harder. We call this quantity downstream work. We introduce a paired replay protocol that forks the same executable state into immediate- handoff and continue-then-handoff branches, separating action-induced work from selection of intrinsically difficult tasks. This protocol supports Dove, a joint distributional model of task outcome and successor work, combined with a resource-aware stopping rule. As a preliminary measurement, we audit 26 public τ -bench result sets comprising 228 tasks, four agent models, and 110,068 trajectory prefixes. Under an evaluator-derived residual-action proxy, 37.3% of trajectories contain at least one work-amplifying action; the task-clustered rate is 29.2% (95% bootstrap CI: 25.8–32.5%). This proxy result motivates, but does not replace, paired successor execution. Our central proposal is that agent evaluation and delegation should optimize correct completed work per system budget rather than autonomous completion alone
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.