REx: Small Language Model Agents with Recursive Execution and Subtree Compression
Abstract
Small language models are attractive for assistants because they can reduce the cost and latency of individual model calls, while local inference can limit exposure of sensitive data. However, long workflows require these models to coordinate dependent steps and retain useful results as execution histories grow. To address these demands, we introduce REx, which combines flexible planning with recursive execution and compression guided by unfinished work. Each planning round selects a batch, an ordered list of one or more subtasks, and assigns each to direct execution or recursive decomposition. Batch length can vary, with updated results guiding replanning after batch completion or subtask failure. To limit context growth, REx selectively replaces completed execution subtrees, the shared records of recursively decomposed subtasks and their nested subtasks, with compact memories. Task goals and any next planned subtask guide which outcomes and evidence to retain for unfinished work. We evaluate REx on Claw-Eval, GAIA, and Agent-Diff with three Gemma-4 models. Under matched models and limits on tool calls, REx outperforms the strongest evaluated baseline in all nine settings, improving task success by up to 19.88 percentage points over standard ReAct. Ablations support flexible planning and subtree compression, while further analysis shows the benefit of supplying task goals during compression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.