acceptodds
Under review as a conference paper at ICLR 2027

REDUCE: Risk-Aware Dual-Level Context Compression for Efficient Terminal Agents

Abstract

The performance of large language model (LLM) agents is increasingly constrained by the accumulation of verbose context. Existing context management methods typically treat rule-based observation filtering and model-based history compression as isolated stages, neglecting the cross-stage risk coupling: Aggressive filtering may remove evidence that downstream summarization cannot recover, whereas conservative filtering passes excessive history to the summarizer, increasing its cost and making it harder to preserve task-relevant details. To bridge this gap, we propose a training-free, self-evolving framework REDUCE, Risk-aware Dual-level Context Compression for Efficient Terminal Agents. It features a risk-aware controller that dynamically governs rule activation, history retention, and compression thresholds by synthesizing online signals, such as context token pressure, summary grounding quality, and rule complaint feedback. Furthermore, it maintains dual timescale persistent knowledge: a global observation rule pool accumulates reusable filtering behaviors, while a statistical memory derives protective priors from historical validation outcomes. Experiments across multiple terminal-agent benchmarks and backbone models show that REDUCE reduces end to end token consumption by 16.1%–63.0% while maintaining or improving task accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.