Generalization in Continuous Flows for Categorical Data
Abstract
Flow matching regresses a velocity field onto a finite training set, and the minimizer of that objective transports noise exactly onto it. Trained models nonetheless generalize, and the cause has been attributed in turn to the stochastic training target, to the sampler, and to the model. These explanations were developed for continuous data such as images, whose distributions are known to concentrate near low-dimensional manifolds. Flow matching has recently been extended to categorical data such as text and protein sequences, whose support is a finite set of points with no such structure, and it is unclear whether and why it generalizes there. We study this question on discrete supports such as permutations and bracket languages. On these supports we derive the optimal denoiser in closed form and show that the flow map it induces is a lookup table from noise to training points. A one-step network regressed onto it is trained on an exact target and never integrates an ODE, yet it still produces valid samples absent from the training set. We measure memorization across training-set size, model depth, training steps and sampling steps, and find that at matched architecture and compute, flow matching generalizes at less data than masked diffusion and less than autoregression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.