acceptodds
Under review as a conference paper at ICLR 2027

Execute What Remains: State-Grounded Skill Execution for Language Agents

Abstract

Language agents can store reusable procedures yet fail to execute them effectively on new tasks. An admissible source action need not advance the current task, and skipping actions with satisfied effects does not determine which stage should proceed next. We introduce State-Grounded Skill Execution (SG), which separates persistent skill knowledge from episode-local submitted progress. SG reconstructs candidates for the remaining acquisition, transformation, and delivery stages, prioritizes direct stage actions, and preserves progress across bounded yields to an incumbent policy. In a prospective comparison on 120 ALFWorld train tasks with no exact-task overlap with prior experiments, all executors use the same skill bank, receiver, and initial budgets. SG improves full-task success over an effect-aware source-order executor (EA) from 30.00% to 46.67%, with 20 SG-only successes and no EA-only success. Mean priced runtime and wall-clock time decrease by 23.16% and 23.18%, respectively, alongside fewer model calls, environment actions, and tokens. On the complete official valid_unseen split, the fixed SG portfolio improves success over the incumbent from 20.90% to 30.60%. The largest gains occur on cleaning tasks that require an intermediate transformation. These results show that controlling a stored procedure's remaining execution can improve both task success and execution efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.