From Preference to Process: Counterfactual Trajectory Reward Modeling for Diffusion RL
Abstract
Terminal preference rewards guide diffusion reinforcement learning, but do not specify how gains should be allocated across interacting denoising transitions. We define counterfactual trajectory credit by selectively replacing stochastic transitions with deterministic updates and replaying their downstream continuations under a fixed policy and indexed noise tape. A frozen preference judge compares the resulting outcomes against a shared deterministic reference, and Shapley values average each intervention's marginal contribution across combinations of retained transitions. To avoid enumerating these outcomes during policy updates, we first adapt clean preference perception to noisy latents, then train the evaluator on offline counterfactual labels to predict step credits and terminal gain from completed trajectory pairs. The predicted credits determine bounded, nonnegative weights on a terminal group advantage. On Stable Diffusion 3.5 Medium, learned allocation improves GenEval from to while retaining the same learned terminal evaluator. Latent initialization reduces step-credit mean absolute error from to , and permutation controls support assigning credits to their corresponding transitions. The results support learned counterfactual credit as useful supervision for allocating terminal feedback across denoising decisions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.