acceptodds
Under review as a conference paper at ICLR 2027

Action2Delta: Decision-Aligned Dynamics Mid-Training for Language Agents

Abstract

Learning action consequences from environmental interactions, typically through next-state prediction, is becoming an important approach to improving language agents’ decision-making. However, high prediction quality can obscure confusion between the consequences of candidate actions and may not reflect how modeling errors affect action selection. We find that conventional generative objectives focus on fitting successor descriptions, without explicitly comparing action consequences from the same starting state or accounting for their effects on decisions. To address this, we propose Action2Delta, a framework for learning from environmental experience to improve decisions. It organizes supervision around action–consequence discrimination and decision relevance through two methods: Contrastive Delta Dynamics (CDD) and Verified Regret Alignment (VRA). CDD executes different actions from the same real environment state and represents their observable consequences as state deltas. It contrasts the likelihood of the same observed delta under different action conditions, strengthening the correspondence between actions and consequences. Under a fixed value evaluator, VRA uses real rollout returns to verify decision gains from replacing predicted successors with real ones. These gains weight dynamics training, while the returns supervise the policy. On ALFWorld and ScienceWorld, static Action2Delta achieves a success rate of 86.57% and a macro-average score of 72.73, respectively, outperforming uniform training with the same real-environment feedback by 1.74 percentage points and 1.80 score points. Evaluation uses only a single actor. Mechanistic analyses further show improved action–consequence discrimination and reduced candidate decision regret.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.