SUCCINCTNESS IS NOT IDENTIFIABILITY: FINITE-EVIDENCE LIMITS FOR TRANSFORMER LENGTH GEN-ERALIZATION
Abstract
Modern Transformers can solve many algorithmic tasks at the lengths observed during training, yet fail on longer inputs. We argue that ffnite endpoint-only evi-dence may fail to identify the intended length-generalizing algorithm. We formal-ize this with the Length-Generalization Identiffability Dimension (LGID), deffned relative to an explicit hypothesis code, observation channel, and long-horizon dis-tribution. For hidden transducer tasks, we separate compact expressibility from identiffability in coded symbolic and programmatic reference classes: endpoint labels can leave compact splice shortcuts, while transition witnesses that cover observable-critical edges remove this shortcut family. Orbit compression remains a conditional corollary under hard equivariance. In controlled hidden-transducer diagnostics, trained Transformers are used as proxy learners for the evidence vari-able identiffed by the theory. Achieved or effective clean transition coverage pre-dicts latent-rule recovery. A train-only readout calibrator gives a behavioral clo-sure diagnostic under ffnal-state label coverage, while endpoint-only and extra-endpoint controls remain near the ffoor. Oracle-free selection and noisy-channel diagnostics show that selector quality is mediated by effective coverage. End-point/readout, transfer, and orbit diagnostics characterize the boundary between latent-rule recovery and full behavioral generalization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.