acceptodds
Under review as a conference paper at ICLR 2027

DomainShuttle: Toward Generalizable Cross-Domain Subject-driven Video Generation

Abstract

Subject-driven text-to-video (S2V) generation has attracted growing interest in academia and industry. An ideal model should faithfully preserve reference appearance for in-domain generation while retaining intrinsic subject features under flexible cross-domain transformations. This generalizability should extend across generation paradigms, from bidirectional fixed-length to streaming autoregressive generation. Existing methods primarily optimize in-domain fidelity by indiscriminately preserving reference appearance, limiting cross-domain flexibility in real-to-fantasy, fantasy-to-real, and real-fantasy interaction scenarios. To address this limitation, we propose DomainShuttle, a novel framework for flexible and high-fidelity video personalization. Specifically, Domain-MoT decouples video and reference processing, and applies Domain-aware AdaLN for domain-specific modeling of reference images. Spatiotemporal Reference Anchor RoPE (STRA-RoPE) temporally anchors reference image tokens and spatially organizes reference tokens with subject-aware offsets, supporting both generation paradigms. For bidirectional training, Cross-Pair Consistent Loss aligns velocity predictions across reference sets to reduce interference from subject-irrelevant features. To further demonstrate its generalizability across generation paradigms, we extend the bidirectional model to streaming autoregressive generation through chunk-wise causal training. Extensive experiments demonstrate that DomainShuttle achieves strong video quality, text controllability, and subject consistency in both in-domain and cross-domain scenarios and on both bidirectional and streaming autoregressive architectures. The anonymous project page with videos is available at https://DomainShuttle.github.io/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.