acceptodds
Under review as a conference paper at ICLR 2027

Native Non-Monotonic Autoregressive Sequence Modeling

Abstract

Autoregressive models generate sequences monotonically, i.e., every sampled token, even if erroneous, becomes a permanent condition for all subsequent steps, and the cost grows with the length of the generation. Non-monotonic autoregressive models address this limitation with a learned erase token, while existing designs roll back one token per command and learn from task-specific augmented traces, so the capability neither scales to long-form generation nor transfers beyond the tasks that supervise it. To this end, we propose N-MARS, **N**ative **N**on-**M**onotonic **A**uto**R**egressive **S**equence models, which introduce a rollback command <UNDO></UNDO> for autoregressive sequence models to roll back tokens as a native generation primitive learned from open-domain text. In particular, we construct a 9.5B on-policy rollback corpus by using model completions of open-domain texts and further employ LLM-as-Judge to annotate the completions, forming the rollback trajectories that allow models to recognize and roll back their own errors. Afterward, we design hybrid continued pre-training (HyCPT), a single stage that unifies continued pre-training and masked supervised fine-tuning under source-specific loss masks, followed by reinforcement learning with outcome rewards that refines when the command fires. Extensive experiments on three long-context generation benchmarks and nine general benchmarks, across four backbones from 3B to 9B parameters in two model families, demonstrate the efficiency and effectiveness of N-MARS.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.