When Does Neural Weight Generation Work? Support, Coordinates, and Evaluation
Abstract
Neural weight generation promises to reuse trained checkpoints, but when does it produce reliable samples rather than a strong best result? We study performance-weighted diffusion on CIFAR-10 CNN and ViT model zoos, varying checkpoint support, generation coordinates, and denoiser capacity. At fixed checkpoint count, replacing one trajectory with three independent runs removes –% of generated accuracy gain above chance under raw-weight coefficient diffusion, despite comparable aggregate denoising loss. Dimensional sweeps expose a normalization artifact that suppresses off-span perturbations, allowing functional outputs despite loss near the zero-predictor baseline. Greater capacity improves higher-dimensional generation, but unequal denoiser fit and exhausted training budgets leave an independent dimensionality penalty unresolved. Whitened single-trajectory span diffusion is the strongest learned custom-zoo generator evaluated, although simple baselines remain competitive and steering generally weakens with parameter-space novelty. On aligned public-zoo supports spanning hundreds of runs, full-latent U-Net diffusion recovers TiltDiff's reported accuracy scale using our own pipeline, with condition-mean Top-1@108 of – across 18 fresh-seed confirmation runs. However, distance-matched noise achieves higher mean accuracy in every run, while exponential weighting improves the upper tail at the cost of lower mean accuracy and yield. These results identify pipeline-specific sensitivities and show why weight-generation evaluation must assess typical-sample quality alongside best-of-samples performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.