acceptodds
Under review as a conference paper at ICLR 2027

Predictions About Unchosen Actions Improve Learning From Agent Trajectories

Abstract

Interactive agents generate trajectories that are later reused for training, yet each decision records an outcome only for the action that was executed. We ask whether the same realized interaction becomes more useful when it also retains pre-action predictions about actions that were available but not chosen. We introduce Candidate logging, which stores the consequence predicted by the trajectory-generating model (the Source) for each candidate action without executing the alternatives. In a controlled latent-rule environment, where an unobserved rule maps mode–action pairs to stochastic consequences, Candidate logging improves held-out consequence prediction from 69.0% to 75.7% (+6.7 percentage points) compared with retaining a prediction only for the executed action. The improvement transfers to autonomous behavior, with no logged predictions or reasoning available at evaluation: relative to the same baseline, task success rises by 5.1 points on ALFWorld valid_unseen (to 52.5%) and by 4.7 points on WebShop (to 39.7%). On both tasks, it also outperforms a baseline that retains the Source’s native pre-action reasoning. Control conditions that keep the logged fields but remove or misalign their action-specific content do not reproduce the controlled gain, which increases as predictions for more unchosen actions are retained and decreases as downstream model capacity and training data grow. These patterns are consistent with explicit predictions reducing the need to recover action–consequence relationships from limited realized experience.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.