acceptodds
Under review as a conference paper at ICLR 2027

Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents

Abstract

Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning enhances task decomposition and tool coordination but also accumulates self-generated text. Over long trajectories, this text can dominate the context and suppress visual evidence, creating textual debt. Our analysis reveals that explicit reasoning steps are superfluous when task-relevant visual evidence is reliably grounded, yet stale hypotheses propagate errors into downstream inference when grounding remains ambiguous. Pruning must therefore remove redundant text without discarding visual evidence. We propose SPARE, a Kullback-Leibler (KL)-guided framework for pruning accumulated reasoning in multimodal tool-use agents. SPARE uses a compact task-state summary as privileged diagnostic context. For each candidate segment, it replays the same model under the original and summary-conditioned contexts. Reverse-KL divergence from on-policy self-distillation (OPSD) then tests whether the summary sufficiently covers the segment without disrupting future reasoning. We further fine-tune the summarizer with supervised fine-tuning (SFT), enabling more compact summaries, broader coverage, and more aggressive pruning. Across five multi-step visual tool-use benchmarks, SPARE achieves the highest average accuracy among pruning methods for all three backbones, improving task accuracy by up to 13.56 percentage points. In the aggregate trade-off analysis between accuracy and pruning, it removes 57.81% of reasoning tokens while surpassing the baseline Pareto frontier. This favorable trade-off shows that reducing textual dominance restores reliance on visual evidence and mitigates over-conditioning on self-generated language.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.