acceptodds
Under review as a conference paper at ICLR 2027

AW-Composer: Compositional Asset and Weather Editing for Driving Videos

Abstract

Compositional editing of driving videos offers a scalable way to create long-tail scenarios by combining asset insertion with weather control. Given a background video, an asset with specified placement, and a retained or newly specified target weather, the goal is to preserve scene structure and asset identity while maintaining spatial and temporal weather consistency. These editing dimensions are coupled: weather transformation changes global appearance and visibility, whereas insertion requires local adaptation to the resulting scene. Achieving this requires coordinated global weather editing and local asset adaptation, yet content-aligned cross-weather supervision is difficult to obtain. We present AW-Composer, a diffusion-based framework that coordinates AW-Weather and AW-Harmony. AW-Weather transforms source videos to requested weather while preserving scene content and also provides paired self-insertion supervision. Trained on these pairs, AW-Harmony uses a single-step, mask-conditioned video DiT with a multi-scale Weather Context Adapter to adapt inserted assets across frames while preserving the scene outside the insertion region. We evaluate weather transformation, paired and open-set harmonization, and specified-weather joint composition. AW-Weather achieves a human score of 93.6/100, compared with 81.2 for VACE. On paired insertion, AW-Harmony reduces FVD by 42.8% relative to the best baseline and achieves the best Weather Consistency on both insertion tracks. On specified-weather joint composition, AW-Composer reduces FVD by 22.5% relative to the strongest harmonization baseline. In the six-way human study, AW-Harmony obtains a Top-1 rate of 49.2%, nearly three times that of the next-best method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.