acceptodds
Under review as a conference paper at ICLR 2027

State Representation Shapes Long-Horizon Agent Learning and Transfer

Abstract

As tool workflows lengthen, agents must select actions over a growing collection of intermediate artifacts. Does supervised training produce behavior that trans- fers beyond the workflow lengths and state representations encountered during training? We investigate this question using 1.7B, 8B, and 27B Qwen models in deterministic workflows that provide the current artifact state and candidate action recipes. We cross training and evaluation state views, evaluate transfer to longer workflows, and measure both action accuracy from correct intermedi- ate states and autonomous task completion. Horizon-matched supervision yields high completion across all three model sizes, but the 1.7B model’s strong short- workflow performance does not ensure transfer to longer workflows. In a repre- sentation crossover with one training seed per view, compact-state training im- proves compact-view accuracy while full-view accuracy declines. The resulting adapter completes all 20 sixteen-action tasks under compact observations but none of the same tasks under full observations. Compaction uses the reference next action to select which artifact contents remain visible, so its effects cannot be at- tributed to reduced state interference alone. In ScienceWorld, compact-state train- ing improves accuracy under both views. Completion also masks execution errors: untuned 27B completes 16 of 20 sixteen-action tasks under full observations de- spite rejected calls in every episode. These results show that supervised gains can remain specific to the training representation, while successful completion can coexist with unreliable execution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.