When Learned Compression Loses Identity: Stage-Resolved Diagnostics for Exact-Entity Conditional Generation
Abstract
Learned compression is effective when task-valid variation can be discarded without identity loss, but can become a misaligned inductive bias when correctness depends on exact, non-substitutable categorical identity. Existing evidence does not determine where exact identity is lost, while comparisons that jointly change representation and corruption cannot attribute failure to either factor. We study personalized multi-day mobility generation as an identity-sensitive testbed, where recent user history conditions future visits over a known location vocabulary, yet models may produce structurally plausible sequences with the wrong locations for that user. An ordered representation–corruption intervention with stage-specific probes localizes identity loss and tests alternative explanations. The diagnosis reveals two distinct mechanisms: vector quantization aliases exact locations before diffusion, whereas continuous autoencoding preserves near-perfect clean reconstruction yet loses prefix-specific identity access during Gaussian latent denoising. Guided by this diagnosis, NADI instantiates the native-mask endpoint through native-space absorbing diffusion and achieves the highest composite objective across all four datasets in the primary evaluation. More broadly, reconstruction fidelity alone does not establish the suitability of a generative representation: for exact, non-substitutable entities, representation and corruption should be judged by whether condition-supplied identity distinctions remain accessible throughout generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.