When Is Tool-Observation Compression Worth Pursuing? Replay Audits Across Agent Interfaces
Abstract
Tool-observation compression is useful only if a transformation applies to the workload, can be decoded from available information, and reduces model-facing cost without damaging behavior. We present a replay audit separating structural support, reuse-weighted mass, paired behavioral leverage and execution integrity. On pinned BFCL tasks, a numeric-array summary saves 41.81% of prompt-plus-completion tokens on eight selected development families but has no qualifying target in 320 held-out reference variants. A typed-subtree state audit gives a different interface comparison: 8.53% of held-out BFCL JSON-object bytes are reconstructible, versus 100% among successful retail and airline reference actions in -Bench. Access to simulator state, however, does not establish decoder access. History-only controls yield little net saving, and neither tokenizer reaches the frozen 5% savings gate in the always-on counterfactual full-template comparison. A supplementary audit of 848 archived histories shows another boundary: structural checks miss same-type referent changes, whereas trusted fingerprints detect the injected changes but can consume the tool-level savings. In a separate 40-position diagnostic, reference encoding reduces exact JSON reconstruction from 80.0% to 37.5% for Qwen and from 52.5% to 5.0% for Llama. The 12-pair natural-task pilot remains inconclusive. Together, these findings show how interface support, decoder access, and cost accounting can invalidate a promising local compression result. The audit supports experiment selection; it does not establish a deployable codec.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.