acceptodds
Under review as a conference paper at ICLR 2027

Fork-NFT: Forked-Credit Negative-Aware Fine-Tuning for Autoregressive Joint Audio–Video Generation

Abstract

Autoregressive joint audio–video generation poses a reinforcement learning problem beyond optimizing complete clips. Each continuation inherits an autoregressive history, while its audio and video streams are tightly coupled: an update may improve one modality while degrading the other or disrupting their synchronization. Consequently, clip-level rewards do not identify which chunks under a shared history should receive video or audio credit. We introduce Fork-NFT, a local reinforcement learning update that branches multiple joint audio–video continuations from a shared prefix and compares them under the same generated history. Modality-specific relative advantages assign video-quality credit to the full continuation after the retained prefix and audio-quality credit to short windows after each fork, while both streams remain jointly generated. To align this local credit with few-step generation, Fork-NFT trains on paired intermediate states recorded from the sampler while re-materializing the causal KV context from each generated prefix under the current policy. We further constrain how cross-modal attention reads the committed history while leaving attention within the active chunk free to adapt. On a 22B eight-step autoregressive model, Fork-NFT improves quality and audio–video consistency over matched baselines on both short- and long-form evaluations, while retaining the same sampler and backbone architecture.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.