acceptodds
Under review as a conference paper at ICLR 2027

SALT: State-Augmented Learning from Trajectories for Language Agents

Abstract

Language agents complete tasks by interacting with users and acting in an environment over multiple turns. At each turn, they choose an action based on the history of user instructions, tool responses, and earlier actions. Standard supervised fine-tuning trains an agent to predict actions from this history, but does not explicitly supervise the interpretation of what the history establishes about the task. We represent that interpretation as a task state that records the agent’s goal, constraints, facts, plan, and open issues. We introduce SALT, State-Augmented Learning from Trajectories, a method that adds task-state supervision to agent training. SALT pairs each action in a trajectory with a task state and trains an agent to generate this state before predicting an action. Across five agentic benchmarks and three Qwen3 model sizes, SALT outperforms standard supervised fine-tuning on the same trajectories in 12 of 15 settings, with an average relative gain of 7%. On -bench, its relative gain at pass^4 is more than double its gain at pass^1. These results suggest that task-state supervision helps agents choose actions that reflect what their interactions have established.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.