acceptodds
Under review as a conference paper at ICLR 2027

ReLaDrift: Recurrent Language Generation via Drifting for Test-Time Scaling

Abstract

The original Drifting Model generates a sample with a single function evaluation, which makes it attractive for fast parallel language generation, but its one-step formulation does not include iterative refinement. We propose Recurrent Language Generation via Drifting (ReLaDrift), which applies one Drifting Model recurrently: each round after the first corrects the draft of the previous round, and training applies the drifting loss to one randomly chosen round without backpropagation through time. On two reasoning tasks with ground-truth answers, Sudoku Hard and GSM8K, the accuracy of ReLaDrift increases with the number of function evaluations (NFE) in the few-step regime (NFE < 32); a single round is already at least as accurate as a single-round Drifting Model trained in the same way, and a few more rounds make it far more accurate. At every NFE from 4 to 32, ReLaDrift outperforms diffusion and flow language models of comparable size trained on the same data, reaching 92.1% on Sudoku Hard at NFE=16 and 29.1% on GSM8K at NFE=4, versus 77.5% and 9.4% for the best of them.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.