acceptodds
Under review as a conference paper at ICLR 2027

When Do Local Score Models Extrapolate Across Size? A Diagnostic Theory and Benchmark

Abstract

Scientific generative modeling often requires size transfer, where models trained on small systems are evaluated on larger ones. While translation-invariant architectures enable this evaluation, we show that architectural locality alone does not guarantee stable size extrapolation. Instead, stable extrapolation depends on how far the Gaussian-smoothed score responds to a perturbation. Even when clean interactions are local, posterior correlations can make the smoothed score respond to distant perturbations. We formalize when this response remains sufficiently local for fixed-receptive-field models to transfer across system sizes, and prove a system-size-independent error bound for generated distributions on fixed local regions. We also introduce Finite-Depth Local Flow (FDLF), a white-box diagnostic benchmark with exact clean scores, densities, and controllable response ranges. Across 2D and 3D FDLF experiments, models transfer stably when the teacher response lies within their receptive fields. A related size-dependent pattern emerges near criticality in the standard 2D Ising model. In a 2D Ising test, the center-score change error of a fixed local model grows with system size near criticality but not away from it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.