acceptodds
Under review as a conference paper at ICLR 2027

Beyond Outcome Feedback: Trajectory-Quality-Guided Experience-based Learning

Abstract

Experience-based learning enables large language model (LLM) agents to improve over time by reusing knowledge from past executions. However, existing methods typically rely on final task outcomes as feedback for experience learning, overlooking the execution process and limiting their ability to evaluate and optimize experience based on trajectory quality. To address this limitation, we propose TQG-Exp, a trajectory-quality-guided experience-based learning framework that introduces process feedback. Specifically, we introduce Trajectory Quality Loss (TQL), which jointly considers task completion and process quality to provide a learning signal beyond final task outcomes. Based on this signal, we employ paired rollouts with and without retrieved experience to guide experience optimization. Representative trajectories are selected based on their outcomes and trajectory quality to extract task-level and task-type-level experiences. Meanwhile, differences in trajectory quality between the paired groups are used to estimate the contribution of retrieved experience and update its utility. The learned utility is then combined with semantic relevance to guide subsequent experience retrieval. Experiments on ALFWorld, WebShop, and ScienceWorld with two backbone LLMs show that TQG-Exp, outperforms existing experience-based methods, improving Avg.@3 over the strongest baseline by 4.50 and 8.55 percentage points on DeepSeek-V4-Flash and Qwen3.5-Flash, respectively.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.