acceptodds
Under review as a conference paper at ICLR 2027

LEARNING FROM DEPLOYMENT FOR CONTINUAL AGENT ADAPTATION

Abstract

Deployment exposes language-model agents to changing tools and workflows, but model providers often receive only interaction records from private client envi- ronments. We study harness adaptation from this deployment experience, without replaying the original environment or sampling new actions during training. We instantiate the setting with two offline updates: experience distillation (OPD) and reinforcement learning (RL) over recorded turns. Distillation extracts procedu- ral lessons, reviews the evidence linking them to recorded operations, and uses an experience-conditioned frozen teacher to supervise the corresponding tokens. The RL update assigns rewards to recorded turns and trains on their complete responses. Across two model sizes and two agent harnesses, both updates improve performance in all four settings on ClawGym dev; distillation also improves three settings on PinchBench. Multi-round experiments with a corrective-and-retention variant of distillation show gains in stationary, cross-agent, and changing-workload settings and examine retention on earlier tasks. These results establish deployment experience as a useful resource for work-agent adaptation, with no access to the client environment during training and no experience prompt at inference time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.