Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents
Abstract
Long-horizon LLM agents must answer queries over histories whose supporting evidence may be localized or dispersed across interaction steps. Lightweight retrieval can miss decisive dependencies, whereas full memory reconstruction incurs unnecessary cost when the retrieved context already suffices. The required processing depth therefore depends on the evidence recovered for a query, not on query complexity alone. We introduce Router-Mem, an evidence-conditioned progressive execution framework that learns whether to stop after a shared retrieval prefix or continue with broader memory processing. To distinguish sufficiency from relevance, we construct supervision by retaining or removing supporting evidence and discarding candidate negatives that remain answerable. We train a lightweight router with binary label supervision and rationale-conditioned representation distillation to make a single-token termination decision without generating rationales at inference time. When continuation is selected, retrieval hits anchor block expansion, parallel memory analysis, and aggregation instead of restarting memory search from scratch. With DeepSeek-V4-Flash at a routing threshold of 0.5, Router-Mem achieves scores of 55.17% on AMA-Bench and 38.77% on BEAM, reducing online memory-processing time by 27.3% and 25.5% relative to full completion, respectively. Code is available at https://anonymous.4open.science/r/Router-Mem-E0A5.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.