π-XLNet: Revisiting XLNet for Text Generation
Abstract
In pursuit of efficient inference and flexible generation, text diffusion architectures have gradually moved closer to XLNet-style arbitrary-order models, motivating us to revisit XLNet's potential for training and generation. We find that, under certain conditions, set diffusion and XLNet are nearly equivalent in architecture. To systematically study this architecture, we propose the π-XLNet framework, which treats the training order distribution as a design variable, and introduce the Mallows distribution to continuously control the degree of training disorder. Experiments show that the additional loss introduced by switching evaluation order decays approximately as a power law as training disorder increases. The training cost for π-XLNet to reach the same natural-order loss as the autoregressive baseline is estimated at approximately 3.1–3.3 times that of the baseline, demonstrating its potential to reduce the training cost of masked text diffusion. Furthermore, we propose the HBD-n decoding algorithm, which uses local future information for stochastic branching and backtracking, translating arbitrary-order conditional prediction ability into generation gains through increased inference computation. The method improves the quality–diversity trade-off over natural-order decoding on a self-trained π-XLNet and raises GSM8K accuracy on a public SetDLM model from 64.67% with official SetDiff decoding to 69.14%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.