Learning from Single-Feedback Deployment Experience
Abstract
The success of LLM agents has attracted growing interest in deploying small-size, self-hosted models to serve as user assistants for everyday work. After deployment, users would like these assistants to improve as they work. In practice, however, tasks arrive sequentially, and users cannot be expected to repeatedly evaluate different attempts at the same task. A common approach is to adapt memories or skills in the agent harness while keeping model parameters fixed. Parameter-based learning provides another promising direction, yet how to effectively train models from such single-feedback deployment experience remains under-explored. In this paper, we study this problem and introduce DEAL, an executable testbed with related task streams in general workplace assistance and AI research. Using DEAL, we find that single-rollout reinforcement learning suffers from an exploration bottleneck, where one evaluated trajectory may leave better behaviors unexplored. Based on this observation, we propose Self-Rewarded Exploration (SRE), which expands additional trajectories from isolated task snapshots and uses the live feedback as an anchor to assign self-rewards. We further provide a theoretical analysis of when the benefit of rollout expansion can outweigh self-reward errors. Experiments on Gemma4-E4B-IT demonstrate the effectiveness of SRE in improving deployment-time performance, generalizing to unseen tasks, and preserving general capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.