Gradient-based Timestep-Partitioned LoRA for Reward Fine-Tuning of Diffusion Models
Abstract
Reward-based fine-tuning enables diffusion models to optimize human preferences and task-specific objectives, but its optimization signals can be highly imbalanced across denoising timesteps. Prior work has identified this issue in policy-gradient-based alignment and addressed it through objective- or scheduler-level correction. We show that concentrated timestep-wise gradients also arise in direct-gradient optimization and that their temporal profile can differ from on-policy reward sensitivity. When a single LoRA adapter is shared across the entire denoising trajectory, concentrated gradients can bias its adaptation toward particular timestep regimes. We propose Gradient-based Timestep-Partitioned LoRA (GTP-LoRA), which partitions the trajectory into contiguous segments and assigns a separate LoRA expert to each segment. Segment boundaries are determined from a reference denoiser-output gradient profile such that each expert receives a comparable cumulative upstream gradient magnitude. This timestep-partitioned design confines gradient updates to their assigned denoising regimes, limiting the disproportionate influence of high-gradient timesteps on a single shared adapter. Experiments across multiple diffusion models, reward objectives, and optimization algorithms show that GTP-LoRA consistently improves reward performance while maintaining competitive sample diversity. The source code is included in the supplementary materials.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.