acceptodds
Under review as a conference paper at ICLR 2027

AsymRefresh: Modality-Asymmetric Refresh for Efficient Joint Audio-Video Generation

Abstract

Feature caching accelerates joint audio–video generation, but joint refreshes tie additional audio computation to video recomputation. We introduce AsymRefresh, a training-free method that makes modality choice part of the refresh decision. Between full joint refreshes, selected steps recompute the native audio pathway against reconstructed video contexts while predicting the video output. A shallow joint prefix and periodically renewed contexts preserve audio access to the evolving video trajectory, with unchanged model weights and samplers. We evaluate JavisDiT++ on all 10,140 JavisBench prompts and both JavisDiT++ and JavisDiT on all 500 T2AV-Compass prompts, covering joint-attention and dual-stream backbones. On JavisBench, AsymRefresh achieves a speedup in a separate timing study with competitive full-benchmark quality against four caching baselines. On JavisDiT, the selected T2AV operating point leads the compared accelerators on four of six primary quality means. Controlled JavisDiT++ studies on 512 prompts and three seeds show improved audio–text alignment over additional joint refreshes at comparable latency under two budgets, with positive multiplicity-adjusted intervals. The six-update gain is also observed at a second prefix depth. Stage and context controls characterize the alignment benefit and accompanying quality and memory trade-offs. These results support modality selection as a useful dimension of inference-budget allocation in joint generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.