acceptodds
Under review as a conference paper at ICLR 2027

Fisher Information and Its Temporal Decomposition for Diffusion Model Alignment

Abstract

Aligning diffusion models with human preferences requires effective use of costly comparison labels. To assess candidate queries by their effects on generation, we study how learning from an answer changes the output distribution and how information about the resulting update is distributed over generation time. We formulate directional Fisher information for updates of a terminal potential, which reweights the model's output distribution. Using Doob's -transform, we connect the sensitivity of the score function to a temporal decomposition of this information. The decomposition characterizes when the terminal value of an update becomes predictable from intermediate states. Terminal Fisher information provides a criterion for query selection, while the decomposition supports analysis of generation. Numerical estimates agree with analytic solutions in a low-dimensional setting. On CIFAR-10, color becomes predictable earlier than car-likeness, and the temporal distribution of update information differs between query selection strategies. In direct preference optimization (DPO) with a fixed diffusion generator, Fisher-based query selection on ImageRewardDB yields lower mean test cross-entropy with 128 labels than random selection with 256 across the evaluated settings. Performance is comparable to D-optimal selection at matched label budgets. When the learned potentials select among shared generated candidates, those learned with Fisher-based query selection yield larger gains than those learned with random selection under the automatic evaluator ImageReward.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.