DFS-GRPO: Reward Guided Tree Search Leads to Provable Improvement in Diffusion Models
Abstract
We propose DFS-GRPO, a reward-guided tree search method that introduces reward-guided multi-tree depth-first search (DFS) sampling to improve diffusion post-training by directing a fixed rollout budget toward high-reward regions while preserving exploration. DFS sampling uses rewards from completed rollouts to select and reuse promising denoising prefixes, progressively directing the sampling groups toward high-reward regions. We combine this sampler with initial noise truncation to strengthen exploitation, while subtree budget allocation and retention of multiple promising branches preserve exploration. An idealized local analysis shows that reward-based selection biases the expected state update toward the gradient of expected reward. Experiments on SD3-Medium and FLUX.1-dev demonstrate that DFS-GRPO improves reward optimization and maintains exploration to guarantee training stability. With matched training rewards, DFS-GRPO outperforms the strongest Flow-GRPO-S3 baseline (ODE–SDE sampling, tree sampling, and noise truncation) across three benchmarks on FLUX.1-dev, achieving relative improvements of on UniGenBench, on GenEval, and on LongText-Bench (English) over this baseline. Under standard inference, the resulting model reaches on OneIG-Bench Alignment, comparable to the newer-generation Z-Image-Turbo and to Imagen-3.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.