Beyond Outcome Feedback: Trajectory-Quality-Guided Experience-based Learning
Abstract
Experience-based learning enables large language model (LLM) agents to improve over time by reusing knowledge from past executions. However, existing methods typically rely on final task outcomes as feedback for experience learning, overlooking the execution process and limiting their ability to evaluate and optimize experience based on trajectory quality. To address this limitation, we propose TQG-Exp, a trajectory-quality-guided experience-based learning framework that introduces process feedback. Specifically, we introduce Trajectory Quality Loss (TQL), which jointly considers task completion and process quality to provide a learning signal beyond final task outcomes. Based on this signal, we employ paired rollouts with and without retrieved experience to guide experience optimization. Representative trajectories are selected based on their outcomes and trajectory quality to extract task-level and task-type-level experiences. Meanwhile, differences in trajectory quality between the paired groups are used to estimate the contribution of retrieved experience and update its utility. The learned utility is then combined with semantic relevance to guide subsequent experience retrieval. Experiments on ALFWorld, WebShop, and ScienceWorld with two backbone LLMs show that TQG-Exp, outperforms existing experience-based methods, improving Avg.@3 over the strongest baseline by 4.50 and 8.55 percentage points on DeepSeek-V4-Flash and Qwen3.5-Flash, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.