Fork-NFT: Forked-Credit Negative-Aware Fine-Tuning for Autoregressive Joint Audio–Video Generation
Abstract
Autoregressive joint audio–video generation poses a reinforcement learning problem beyond optimizing complete clips. Each continuation inherits an autoregressive history, while its audio and video streams are tightly coupled: an update may improve one modality while degrading the other or disrupting their synchronization. Consequently, clip-level rewards do not identify which chunks under a shared history should receive video or audio credit. We introduce Fork-NFT, a local reinforcement learning update that branches multiple joint audio–video continuations from a shared prefix and compares them under the same generated history. Modality-specific relative advantages assign video-quality credit to the full continuation after the retained prefix and audio-quality credit to short windows after each fork, while both streams remain jointly generated. To align this local credit with few-step generation, Fork-NFT trains on paired intermediate states recorded from the sampler while re-materializing the causal KV context from each generated prefix under the current policy. We further constrain how cross-modal attention reads the committed history while leaving attention within the active chunk free to adapt. On a 22B eight-step autoregressive model, Fork-NFT improves quality and audio–video consistency over matched baselines on both short- and long-form evaluations, while retaining the same sampler and backbone architecture.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.