acceptodds
Under review as a conference paper at ICLR 2027

Learning Reusable Partial-Order Skills from Agent Traces

Abstract

Agent traces are sequential, but the procedures they record may allow alternative action orders, repetition, and repair. We introduce Reusable Partial-Order Skills (RPOS), a probabilistic model that jointly infers skill boundaries, assignments, and shared precedence constraints from action-labelled traces. The action dictionary, number of skills, are specified in advance; skill decompositions and partial orders are inferred. Each skill instance tracks the validity of action outputs, modelling repetition and repair without adding cycles to its shared graph. Semi-Markov dynamic programming sums exactly over admissible segmentations and skill assignments to evaluate structural proposals. Experiments on synthetic data show improved recovery and prediction in the completed runs. On coverage-controlled TaskBench simulations, RPOS predicts unseen valid action orderings better than the tested total-order and sequence models. On -Retail, RPOS learns an explicit library of partial-order skills and transitions between them, making within-skill precedence relations and between-skill transitions transparent. Finally, on 120 paired synthetic Retail tasks, an executor combining the learned library before evaluation reduces online token use by 86.97% relative to a stepwise agent while meeting a prespecified non-inferiority margin for strict task success.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.