acceptodds
Under review as a conference paper at ICLR 2027

Inside the Scaling of Riemannian Flow Language Models: Geometry, Generation, and Conditioning

Abstract

Riemannian flow language models have not been studied at practical scale. We train SwaFlow, a suite of hyperspherical flow language models with up to 3B parameters on 2T tokens. As training progresses, token embeddings become less isotropic, shrinking the range of noise levels from which local geometry alone recovers a token. Yet the models recover tokens from increasingly higher noise levels using context, and the gap between these regimes grows with scale. Consistent with this trend, held-out loss and unconditional generation improve with model size and training budget, but conditioning on clean context does not: multiple-choice accuracy stays at chance at every size and budget. We trace this failure to the training distribution rather than the geometry: standard training assigns the same noise level to all positions and never exposes the model to mixed clean–corrupted contexts. Prior work under a different geometry has shown that clean-prefix corruption improves conditional generation; for Riemannian flow language models, scaling alone does not remove this dependence on partially clean training examples. Introducing them only during the final 100B training tokens lifts conditioning off chance at every size, with gains that grow with model size and training budget and carry over to task finetuning on GSM8K. We will release all intermediate checkpoints, optimizer states, and exact data order upon acceptance, enabling these scaling trajectories to be reproduced, branched, and re-evaluated.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.