acceptodds
Under review as a conference paper at ICLR 2027

ReCompact: Learning to Act and Compact from Graph-Guided Trajectory Supervision

Abstract

Long-horizon software agents compress interaction histories to operate within a finite context window. However, a mistaken edit is a poor imitation target, even though it remains part of the working tree after a history reset. We introduce **ReCompact**, which separates **action weighting** from **state retention** through graph-guided trajectory supervision. A supervised neural graph module assigns contribution labels to weight action targets and prioritize key and supporting information in compact training targets and online compaction. Two complementary training views jointly supervise compact generation and continuation after a history reset. Remarkably, on the same **500 SWE-bench Verified tasks**, ReCompact resolves **49.8%** of tasks, substantially outperforming **standard supervised fine-tuning (29.0%)**, **step-aware training (43.4%)**, and **compact-only training (39.4%)**. While achieving a **62.4% median input reduction**, comparable to the **61.2%** reduction of compact-only baselines, ReCompact nearly halves repetitions of previously failed commands, reducing them from **7.64 to 4.01 per 100 post-compaction tool actions**. Ablations further confirm the importance of each component: disabling **online compaction and reset** reduces task resolution by **5.8 percentage points**, while removing **continuation training** or **retention-label guidance** causes drops of **4.6** and **3.4 percentage points**, respectively. ReCompact is also more inference-efficient, using **10.1% fewer total inference tokens per task** than compact-only alternatives.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.