Fisher Information and Its Temporal Decomposition for Diffusion Model Alignment
Abstract
Aligning diffusion models with human preferences requires effective use of costly comparison labels. To assess candidate queries by their effects on generation, we study how learning from an answer changes the output distribution and how information about the resulting update is distributed over generation time. We formulate directional Fisher information for updates of a terminal potential, which reweights the model's output distribution. Using Doob's -transform, we connect the sensitivity of the score function to a temporal decomposition of this information. The decomposition characterizes when the terminal value of an update becomes predictable from intermediate states. Terminal Fisher information provides a criterion for query selection, while the decomposition supports analysis of generation. Numerical estimates agree with analytic solutions in a low-dimensional setting. On CIFAR-10, color becomes predictable earlier than car-likeness, and the temporal distribution of update information differs between query selection strategies. In direct preference optimization (DPO) with a fixed diffusion generator, Fisher-based query selection on ImageRewardDB yields lower mean test cross-entropy with 128 labels than random selection with 256 across the evaluated settings. Performance is comparable to D-optimal selection at matched label budgets. When the learned potentials select among shared generated candidates, those learned with Fisher-based query selection yield larger gains than those learned with random selection under the automatic evaluator ImageReward.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.