State Representation Shapes Long-Horizon Agent Learning and Transfer
Abstract
As tool workflows lengthen, agents must select actions over a growing collection of intermediate artifacts. Does supervised training produce behavior that trans- fers beyond the workflow lengths and state representations encountered during training? We investigate this question using 1.7B, 8B, and 27B Qwen models in deterministic workflows that provide the current artifact state and candidate action recipes. We cross training and evaluation state views, evaluate transfer to longer workflows, and measure both action accuracy from correct intermedi- ate states and autonomous task completion. Horizon-matched supervision yields high completion across all three model sizes, but the 1.7B model’s strong short- workflow performance does not ensure transfer to longer workflows. In a repre- sentation crossover with one training seed per view, compact-state training im- proves compact-view accuracy while full-view accuracy declines. The resulting adapter completes all 20 sixteen-action tasks under compact observations but none of the same tasks under full observations. Compaction uses the reference next action to select which artifact contents remain visible, so its effects cannot be at- tributed to reduced state interference alone. In ScienceWorld, compact-state train- ing improves accuracy under both views. Completion also masks execution errors: untuned 27B completes 16 of 20 sixteen-action tasks under full observations de- spite rejected calls in every episode. These results show that supervised gains can remain specific to the training representation, while successful completion can coexist with unreliable execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.