LEARNING FROM DEPLOYMENT FOR CONTINUAL AGENT ADAPTATION
Abstract
Deployment exposes language-model agents to changing tools and workflows, but model providers often receive only interaction records from private client envi- ronments. We study harness adaptation from this deployment experience, without replaying the original environment or sampling new actions during training. We instantiate the setting with two offline updates: experience distillation (OPD) and reinforcement learning (RL) over recorded turns. Distillation extracts procedu- ral lessons, reviews the evidence linking them to recorded operations, and uses an experience-conditioned frozen teacher to supervise the corresponding tokens. The RL update assigns rewards to recorded turns and trains on their complete responses. Across two model sizes and two agent harnesses, both updates improve performance in all four settings on ClawGym dev; distillation also improves three settings on PinchBench. Multi-round experiments with a corrective-and-retention variant of distillation show gains in stationary, cross-agent, and changing-workload settings and examine retention on earlier tasks. These results establish deployment experience as a useful resource for work-agent adaptation, with no access to the client environment during training and no experience prompt at inference time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.