acceptodds
Under review as a conference paper at ICLR 2027

Words, not Walks: An Algebraic Theory of Tool-use Trajectory Reduction for Agent Distillation

Abstract

A language agent that solves a task by calling tools produces a trajectory: a sequence of actions that transforms an environment state. We argue that this process is most naturally described not as a walk on a graph but as a word in a transformation monoid: tools are generators, dependencies are non-commutativity, and an efficient solution is a geodesic—a shortest generator product. This view exposes two failure modes that plague behavioral cloning of agents: state-relative identity sub-words (no-ops and backtracking loops that fix the current state) and non-commutative ordering errors. We introduce Algebraic Trajectory Reduction (ATR), a sound, model-free procedure that rewrites a teacher trajectory to its empirical geodesic before it is used for supervised fine-tuning. ATR unifies loop elimination and observed shared-state shortcuts as shortest-path extraction over the observed transition graph, with provable soundness under deterministic tools and exact state identifiers. We further give a sufficient boundary: on strictly monotone response integration, ATR coincides with deduplication by construction, whereas state revisits in sequential tool-use can make dedup unsound. Both regimes appear in real data. On 44,730 GPT-4/DFSDT TOOLBENCH trajectories, 27.9% of API calls fail (no-information updates) and ATR removes 39.4% of calls, exactly matching the strongest baseline as predicted. On 336 sequential AGENTINSTRUCT/ALFWorld trajectories, ATR removes twice as much as dedup (24.4% vs. 12.0%); under the observation proxy, offline replay fails on 1.5% of ATR paths versus 25.3% for dedup, and the former 1.5% is already present in unreduced data. A conservative certified variant removes only 1.3%. A 3B Qwen student trained on the proxy-reduced supervision achieves closed-loop success comparable to RAW/B2 within seed variation (9.5% vs. 9.2%/10.0%) using 24.4% fewer action targets. In a controlled transformation-monoid environment where teacher verbosity is a manipulated variable, cloning identical policies on ATR-reduced data raises closed-loop task success from 71.4% to 97.2% (matching an oracle-geodesic skyline of 97.8%), cuts non-commutativity errors from 40.9% to 0.7%, and uses 46% fewer supervision transitions; ATR is invariant to how verbose the teacher is, whereas naive cloning degrades monotonically. Our results give a precise algebraic account of what makes an agent trajectory worth imitating.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.