From Trajectory Memory to Action Selection: Trajectory-Context Action Credit for Language Agents
Abstract
Language agents increasingly reuse past interaction trajectories as examples, memories, abstractions, and decision signals. However, retrieving a relevant trajectory does not directly tell an agent which action it should prefer at the current Decision State. A trajectory contains temporally extended experience, while action selection requires choosing among the Candidate Actions available now. We formalize this missing link as Trajectory-Context Action Credit, a candidate-level signal that makes the contribution of trajectory memory to action selection explicit. We introduce TrajCredit, a training-free instantiation that converts Retrieved Slices into candidate-level credit using Local State-Action Support, Procedural-Stage Consistency, and Outcome-Calibrated Evidence, and combines the resulting credit with an existing Base Selector. Across controlled candidate-action ranking, retrieval and memory controls, credit-to-progress alignment, and bounded environment intervention, TrajCredit consistently improves action ranking, provides information beyond raw retrieval support, aligns with independently measured short-horizon value, and improves executed behavior. These results support Trajectory-Context Action Credit as a useful interface between reusable agent experience and candidate-action selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.