acceptodds
Under review as a conference paper at ICLR 2027

FoldAct: Stable and Efficient Training for Context-Folding Agents

Abstract

Long-horizon reinforcement learning (RL) for large language model agents operates under fixed context windows while interaction histories keep growing. Context folding summarizes this history, but end-to-end folding changes the training problem because the same policy writes summaries and later acts from them. Summary tokens are policy-written memory—tokens that become future visible state—so a unified token-level objective creates credit dilution, summary-induced replay mismatch, and expensive per-turn replay. We introduce FoldAct, an RL training method for policy-written memory in long-horizon agents. FoldAct combines separated summary/action objectives, a consistency loss that stabilizes behavior after context folding, and selective segment training. At inference, the same policy writes summaries and produces subsequent actions, without a teacher summarizer, external memory model, or change to the deployed architecture. Our controlled claims are feasibility, update-side cost, and stability: FoldAct makes folded training feasible with a 5.19× speedup and stabilizes response length.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.