Off-Diagonal Atlas Transport
Abstract
Next-token prediction determines what a language model must predict, but does not uniquely determine how task-relevant information is organized across token positions. Temporal representation objectives can shape this organization. However, matching the outputs of a learned projection across positions admits an easy shortcut: a persistent output direction can satisfy the objective without tracking hidden-state changes. We introduce Off-Diagonal Atlas Transport (OAT), a training-time regularizer that maps each hidden state through multiple learned projections and rewards temporal compatibility only across different projection indices. Because off-diagonal matching can itself become trivial when projections of the same state become nearly identical, OAT also suppresses positive overlap among them. Across three instruction-tuned language-model families and three low-data fine-tuning tasks, OAT improves mean accuracy over next-token prediction in all nine backbone–task combinations, with its clearest gains on mathematical reasoning. Ablations show that allowing same-index temporal matches or removing cross-projection redundancy suppression makes the auxiliary objective easy to satisfy. Yet the downstream benefit largely disappears under these changes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.