Measuring Order Learnability via Prequential Codelength
Abstract
Does sequential data favor left-to-right prediction? Although entropy is invariant to factorization order, a finite learner may acquire some factorizations more efficiently than others. We measure order-learnability using prequential codelength, which accumulates the log-loss incurred when a model predicts each batch before training on it. After synthetic calibration, we compare 21 fixed orders across 13 corpora under a controlled Transformer training protocol. Our experiments reveal four findings. First, left-to-right is rarely uniquely favored: it is significantly better than every tested alternative only on mathematical chain-of-thought traces. Second, text is substantially more sensitive to local shuffling than to full reversal, suggesting that preserving local adjacency matters more than preserving forward direction for this learner. Third, the differences between orders change with training: both additional data and repeated passes narrow the relative gap on text but widen it on two protein sources. Fourth, frequency-sorted orders achieve substantially lower prediction loss on English Wikipedia while supplying additional information about the content. Our framework turns intuitions about a corpus's preferred prediction order into quantitative comparisons of learning efficiency under a fixed learner.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.