acceptodds
Under review as a conference paper at ICLR 2027

LongStraw: Memory-Constrained Policy Updates with Million-Token Prefixes

Abstract

Policy training on long histories remains memory intensive even with low-rank adapters, because sequence state and activations grow with the input. We introduce LongStraw, a framework for memory-constrained policy updates on complete long inputs. It couples a compact, versioned prefix-state data structure with a capture–replay–recapture algorithm. The structure organizes layer-specific continuation tensors with their ownership, positions and gradient boundary, sharing prompt state across responses while keeping their extensions private. Capture constructs this state without autograd; block-wise reverse replay differentiates responses and an optional short prompt tail; recapture refreshes the state after each optimizer step. This bounds live activation memory while preserving access to the complete input. The update optimizes a conditional objective that omits gradients through construction of the captured prefix. On one H200, LongStraw completes million-token Qwen3.6-27B updates where native CPU-offloaded autograd of the same objective runs out of memory. On a fixed million-token question-answering benchmark, exact-match accuracy improves by 2.03 percentage points across paired training seeds. Experiments show learning gains on complete inputs up to 4M tokens, and distributed GLM-5.2 experiments demonstrate online policy updates. These results show how jointly organizing prefix state, backward computation and state refresh makes policy training on complete long histories feasible under memory constraints.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.