Latent Parity-Check Tokens: Making Transformer Representations Internally Verifiable
Abstract
Transformer representations are typically treated as opaque intermediate states. When latent features become corrupted or inconsistent, subsequent layers continue to process them without explicitly detecting the failure. We introduce Latent Parity-Check Tokens (LPCT), a structured redundancy mechanism that makes Transformer representations explicitly verifiable. Inspired by parity-check codes, LPCT augments semantic tokens with learned check tokens connected through a structured topology. These relationships define testable consistency constraints whose violations produce a latent syndrome that can detect and localize unreliable latent states while optionally supporting reconstruction. Across controlled intermediate-token corruptions, LPCT maintains competitive predictive robustness while its syndrome nearly perfectly detects latent corruption (AUROC ) and increasingly localizes corrupted tokens as erasure severity increases (0.565-0.948 AUROC). These results suggest that structured redundancy can provide explicit internal consistency checking in Transformer representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.