acceptodds
Under review as a conference paper at ICLR 2027

DFS-GRPO: Reward Guided Tree Search Leads to Provable Improvement in Diffusion Models

Abstract

We propose DFS-GRPO, a reward-guided tree search method that introduces reward-guided multi-tree depth-first search (DFS) sampling to improve diffusion post-training by directing a fixed rollout budget toward high-reward regions while preserving exploration. DFS sampling uses rewards from completed rollouts to select and reuse promising denoising prefixes, progressively directing the sampling groups toward high-reward regions. We combine this sampler with initial noise truncation to strengthen exploitation, while subtree budget allocation and retention of multiple promising branches preserve exploration. An idealized local analysis shows that reward-based selection biases the expected state update toward the gradient of expected reward. Experiments on SD3-Medium and FLUX.1-dev demonstrate that DFS-GRPO improves reward optimization and maintains exploration to guarantee training stability. With matched training rewards, DFS-GRPO outperforms the strongest Flow-GRPO-S3 baseline (ODE–SDE sampling, tree sampling, and noise truncation) across three benchmarks on FLUX.1-dev, achieving relative improvements of on UniGenBench, on GenEval, and on LongText-Bench (English) over this baseline. Under standard inference, the resulting model reaches on OneIG-Bench Alignment, comparable to the newer-generation Z-Image-Turbo and to Imagen-3.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.