Never Optimize Your Memory-Compression Harness Under Training–Inference Mismatch
Abstract
Production agent harnesses, including Claude Code and Qwen-Agent, compress an agent's context during rollout to make long-horizon interaction tractable. Training policies under such compression, however, introduces a fundamental conditioning problem: each compression-induced eviction branches the effective interaction history, making the object presented to the learning objective a tree rather than a single trajectory sequence. Existing learning pipelines typically linearize this tree in one of two ways: (1) retaining only the rightmost root-to-leaf path, which causes time-travel leakage, or (2) replaying the full depth-first traversal, which induces a train–inference mismatch. Thus, we introduce two exact conditioning-level corrections: (1) LogitTree, a segmented K-forward traversal of the trajectory tree, and (2) an equivalent packed Logit attention mask. LogitTree requires K+1 backward passes for every root-to-leaf, whereas the Logit formulation requires both a custom masked-attention kernel and white-box access to harness eviction records. To avoid these costs, we further propose SDCC (Self-Distillation for Conditioning Consistency), a training-friendly variational relaxation requiring only a single backward pass and multiple parallel forward passes. At each eviction junction, SDCC minimizes the forward KL divergence from a stop-gradient teacher distribution conditioned on the reconstructed pre-eviction prefix to the compressed-context student policy. By Pinsker's inequality, a residual per-junction KL divergence of ε_KL yields an O(√ε_KL) bound on the resulting train–deployment behavioral gap in total variation. In black-box harnesses, SDCC uses captured per-turn request payloads and requires no span-level eviction logs. Experiments across three white-box editors, two black-box harnesses, and seven web-search benchmarks reveal substantial conditioning drift under naive replay. LogitTree and Logit masking recover the no-compression numerical baseline, while SDCC substantially narrows the gap with a modest computational budget. On the WideSearch benchmark, our LogitTree method improves the Item F1 of Qwen3.7-Air (125B-A6B) from 67.32 to 74.03, surpassing all reference models in our evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.