Field Constrained Bargaining DiffusionNFT
Abstract
Diffusion alignment increasingly requires optimizing several criteria at once, such as visual quality, prompt fidelity, text rendering, and compositional correctness. Existing approaches commonly aggregate these criteria with fixed reward weights or balance their stochastic parameter gradients. Both views are incomplete for Diffusion Negative-aware FineTuning (DiffusionNFT), whose primitive learning signal is a scalar optimality that reweights the current clean-image distribution. A reward vector therefore has to be converted into a scalar target, and a target that appears attractive in reward space may still require a large modification of the forward velocity field. We formulate multi-reward DiffusionNFT as a constrained target-selection problem. For the linear optimality family , we show that the exact distribution-level reward gain is , where is the covariance of the centered rewards. We further derive an NFT-native field Gram matrix for which the population-optimal velocity displacement is exactly . These two geometries lead to Field-Constrained BargainingNFT (FCB-NFT), a small convex program that maximizes Nash reward surplus under a velocity-field displacement budget and valid-optimality constraints. We also provide a finite-sample conservative estimator of that avoids learning a conditional field probe. On two four-reward suites, FCB-NFT alone improves every reward, with a worst normalized gain of while baselines stay imbalanced or collapse. Ablations trace this gain to the structure of beyond reward covariance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.