acceptodds
Under review as a conference paper at ICLR 2027

What Makes a Search Space Effective for Evolution Strategies in Diffusion Fine-Tuning?

Abstract

Evolution strategies (ES) fine-tune generative models from reward evaluations alone, yet successful applications search very different spaces, from all weights of language models to about a hundred scaling factors of diffusion Transformers. Classical analyses of zeroth-order optimization favor small search spaces, but a smaller space also restricts what fine-tuning can change, and spaces of different sizes usually differ in how their parameters act on the network. It therefore remains unclear what makes a search space effective for ES on diffusion models. We separate these two factors in more than one hundred ES runs on SD3.5-Medium, FLUX.1-dev and Wan2.1-T2V, comparing search spaces at matched perturbation energy. Effectiveness depends more on how each parameter acts than on how many parameters there are. On SD3.5-Medium, growing residual scaling from 24 to 36,864 tunable parameters leaves the reward improvement unchanged, even as the agreement between independent update estimates drops by more than an order of magnitude. At 24 parameters, by contrast, scaling whole residual branches improves the reward by 1.68%, dense random mixtures of channels by 0.75% and random channel subsets by only 0.12%. ES search spaces can therefore be large, provided each parameter rescales a whole architectural unit. The improvement generalizes to held-out prompts and reproduces on other backbones, but unoptimized metrics stay flat or decline and optimizing a compositional objective fails, so the reward must encode every property that fine-tuning is meant to improve.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.