PRECOMPUTED DECODER JACOBIANS FOR FEW-STEP AUDIO GUIDANCE
Abstract
Training-free guidance steers a pretrained generator with the gradient of a score on its output. In latent audio generators this score sits behind a decoder, so every guidance step pays for differentiating it, and the common remedy trains a latent head that imitates the decoder's gradient. We observe that for attributes such as loudness the decoder's Jacobian changes little from one latent state to another, and propose Mean-Jacobian guidance (MaJa), which replaces it with a single matrix: the Jacobian averaged over a handful of training states, computed once and reused for every state, prompt and step. Paired with a head that predicts only the attribute value, this takes the decoder out of the sampling loop and makes guidance an order of magnitude cheaper. On few-step samplers the shortcut comes without a loss of accuracy in almost every setting. Although its direction has a cosine of only 0.60 to 0.91 to the decoder gradient, MaJa is level with or ahead of differentiating the decoder in fourteen of sixteen settings across two generators and five attributes, and reduces target error by 31% on average relative to learned latent guidance heads. Closer imitations of the gradient do not help: a head that reproduces it at cosine up to 0.99 controls no better, and the exact gradient through the remaining sampler steps does worse. On these samplers, agreement with the decoder gradient is not what makes a guidance direction effective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.