DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation
Nan Chen, Yiyang Cai, Rongchang Xie, Junwen Pan, Cheng Chen, Weinan Jia, Zhuowei Chen, Wen Zhou, Zhenbang Sun, Wenhan Luo
arXiv:2606.26058 · 67▲ · cs.CV
View on arXiv →Premium readers get an interactive explainer for this paper — a figure you can poke at, not just read.
What it is
DomainShuttle is a subject-driven text-to-video system built on the Wan2.1/2.2 DiT video models that adds three components: Domain-MoT (separate transformer branches for reference images versus video plus a domain-aware AdaLN modulation), Video-Reference DualRoPE (placing reference tokens in a separate positional encoding space from video tokens), and a Cross-Pair Consistent Loss that aligns two different reference sets of the same subject. The goal is to preserve a reference subject's intrinsic features while still allowing style, domain, and semantic changes driven by the text prompt (cross-domain generation), rather than only copying the reference (in-domain).
Why it matters
For anyone building video personalization tools (advertising, creative design, AI filmmaking), existing methods tend to copy-paste the reference subject and resist stylistic transformation, so you cannot easily put a real person into a watercolor or 3D-animation scene while keeping their identity. This method aims to keep identity fidelity and still follow style instructions, which is the practical difference between a rigid template filler and a flexible editor.
Practical takeaway
Watch for this in Wan-based video generation tooling: it is fine-tuned on top of open-source Wan2.1/2.2-14B, so if you already use those models, the approach (decoupled reference branch, separate RoPE space for reference tokens, consistency loss across reference sets) is a recipe you could adopt for subject-driven generation that needs style transfer, not just identity locking.
Key result
On the authors' own 220-sample test set (110 in-domain, 110 cross-domain), the Wan2.2 version reports a Cross-Domain Score of 0.861 versus 0.725 for the best baseline (Kling 1.6), an 18.7% relative improvement. Note the evaluation set is self-constructed and modest in size, and several cross-domain metrics rely on the authors' own MLLM-based scoring (GPT-5.2, Qwen3-VL) and image-edit-based CLIP pipelines rather than an established public benchmark.
Subscribe
Get the next issue.
Free. One email a week. Unsubscribe any time: no account, no dark patterns.