acceptodds
Under review as a conference paper at ICLR 2027

Effects of Token-level Loss Reweighting on Long-Range Recall in Chunked Training of Mamba-2

Abstract

Truncated backpropagation through time trains a stateful sequence model by carrying the state across chunks while stopping gradients at chunk boundaries, which biases learning toward short-range dependencies. A standard approach is to use a longer gradient window. In this study, we investigate whether extending the gradient path is necessary for improving long-range recall. Our hypothesis is that, with the state carried forward, which tokens the objective emphasizes can matter more for long-range recall than how far the gradient reaches. We test this hypothesis in continued training of Mamba-2 (1.3B and 370M) on PG-19, varying two choices independently: (1) whether the loss is the standard cross-entropy or is reweighted toward rare tokens absent from the recent context, and (2) whether gradients stop at chunk boundaries or flow through the whole sequence. We deliberately use a simple reweighting scheme as a probe of the hypothesis: a deterministic mask on token ids with a single weight set by a static calibration rule. With the gradient window fixed at one chunk, reweighting raises variable-tracking accuracy at the 8K training length from 18.3 to 33.1 and needle retrieval at four times the training length from 67.7 to 94.0. At 1.3B and the 8K training length, full backpropagation gives no statistically significant improvement under the standard objective. The gains are task-dependent and trade off against perplexity and accuracy on an aggregation task (frequent-word extraction). In this setting, long-range recall improved by changing what the objective emphasizes, without extending the gradient path.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.