The Context of Theseus: Learning Selective Discrete Recurrence in Autoregressive Transformers
Abstract
Evolving programs, proofs, and plans often requires iteratively expanding and revising intermediate structures. Standard autoregressive decoding, however, restricts sequence evolution to monotonic prefix extension, binding logical sequence order to causal generation order and indefinitely accumulating superseded drafts in context. To address this problem, we introduce Theseus-AR, which fine-tunes pretrained autoregressive Transformers for selective discrete recurrence: iteratively evolving an explicit sequence state through in-place transitions that selectively retain, revise, or retire predecessor content to condition subsequent steps. Through RoPE manipulation and active key–value reorganization, Theseus-AR maintains a topologically coherent representation across transitions without re-encoding successor states from scratch. To train this recurrence efficiently, we develop Replay Attention, which compiles dynamic topology views along the full sequence of state transitions into parallel replay units that reuse shared representations, restoring token-level parallel training with exact-arithmetic equivalence to online execution while reducing visible query–key attention pairs from quadratic in transcript length to workload-bounded . Across programmatic testbeds, Theseus-AR matches the Transition Trace validation-error level with up to fewer supervised target tokens in surgical in-place rewrites, reaches comparable validation error with 54%–87% fewer visible attention pairs in complex regimes, improves deep-horizon transition validity under Stepwise Transition Evaluation, and reduces peak active context tokens at inference by 59%–88%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.