Environment-Free Learning from Agent Trajectories with Bottleneck Upsampling
Abstract
Language agents are increasingly deployed in complex environments, and every deployment leaves behind trajectories. Rebuilding the environment for post-training, as reinforcement learning requires, is often impractical: it may be proprietary, involve real users, or be too expensive to run again. Environment-free learning, which improves an agent from these trajectories alone, is therefore an important problem. A natural approach is to relabel the trajectories: at every step, a teacher shows what should have been done, and the agent is trained to follow. In context distillation (CD) and on-policy context distillation (OPCD), the teacher is the same model conditioned on experience summarized from the trajectories. We find that this recipe works well on some benchmarks but delivers limited gains on others, and we trace the shortfall to the trajectories. The deployed agent that collected them is weak, so few episodes reach the later stages of a task, and the training data hold few steps from them. For example, on a shopping task, few episodes reach a product page, so the trained agent does not learn what to do there, such as choosing options and placing the order. We call the first stage that few episodes reach the bottleneck. We propose CDUP (Context Distillation with bottleneck UPsampling), a simple and effective recipe that has an LM agent segment the trajectories into task stages, locates the bottleneck from the share of episodes reaching each stage, and upsamples the steps from it onward at a fixed ratio. On WebShop, CDUP improves success over CD by 16.4 percentage points with Llama-3.2-3B and 3.1 with Qwen2.5-3B, and on ScienceWorld by 16.9 and 9.9. Bottleneck upsampling also raises WebShop success under OPCD and with an external expert teacher. Code is available at https://anonymous.4open.science/r/cdup.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.