acceptodds
Under review as a conference paper at ICLR 2027

When Does Neural Weight Generation Work? Support, Coordinates, and Evaluation

Abstract

Neural weight generation promises to reuse trained checkpoints, but when does it produce reliable samples rather than a strong best result? We study performance-weighted diffusion on CIFAR-10 CNN and ViT model zoos, varying checkpoint support, generation coordinates, and denoiser capacity. At fixed checkpoint count, replacing one trajectory with three independent runs removes –% of generated accuracy gain above chance under raw-weight coefficient diffusion, despite comparable aggregate denoising loss. Dimensional sweeps expose a normalization artifact that suppresses off-span perturbations, allowing functional outputs despite loss near the zero-predictor baseline. Greater capacity improves higher-dimensional generation, but unequal denoiser fit and exhausted training budgets leave an independent dimensionality penalty unresolved. Whitened single-trajectory span diffusion is the strongest learned custom-zoo generator evaluated, although simple baselines remain competitive and steering generally weakens with parameter-space novelty. On aligned public-zoo supports spanning hundreds of runs, full-latent U-Net diffusion recovers TiltDiff's reported accuracy scale using our own pipeline, with condition-mean Top-1@108 of – across 18 fresh-seed confirmation runs. However, distance-matched noise achieves higher mean accuracy in every run, while exponential weighting improves the upper tail at the cost of lower mean accuracy and yield. These results identify pipeline-specific sensitivities and show why weight-generation evaluation must assess typical-sample quality alongside best-of-samples performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.