Rewind the Transformer: Input-Anchored Computational Replay
Abstract
Transformer states have continuous coordinates, but the states a node can take are generated by discrete admissible sources. Reversibility is then a question about this realizable set: which source distinctions does an observation preserve? When the input's identity survives, recovering it once lets one deterministic forward pass replay every declared state without inverting any layer. For a lossy code fixed before recovery, a hybrid decoder built on SipIt and first-order token ranking makes the route executable: cheap search proposes candidates, only the observation's own rule accepts them, and ordinary exhaustion falls back once to SipIt. Behind a frozen 8-bit-per-coordinate interface on openai-community/gpt2, both a SipIt-based decoder and the hybrid recovered the 8-token input of each of 8 separately selected documents and reproduced all 14 declared nodes bitwise, whereas continuing from the dequantized state shifted the next-token distribution on every document (top prediction still agreed on all 8). In a fitted-proposal extension on 16 fixed news inputs, an affine proposer completed all 16 in both repetitions, versus 15 for non-speculative continuous recovery under a 120 s decode cap, with a median 3.5× request-time speed-up on the 15 common completions. Diagnostics at known prefixes show these distinctions coming apart for geometry, read-outs, partial views and a known edit. Discreteness alone does not guarantee distinguishability and a match is not a uniqueness certificate; these results establish a small-scale workflow, not a generally superior codec or a proof of unique recoverability over all inputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.