Fig.1

From Issue #8 · 2026-06-29

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation

Nan Chen, Yiyang Cai, Rongchang Xie, Junwen Pan, Cheng Chen, Weinan Jia, Zhuowei Chen, Wen Zhou, Zhenbang Sun, Wenhan Luo

arXiv:2606.26058 · 67▲ · cs.CV

View on arXiv →

Premium readers get an interactive explainer for this paper — a figure you can poke at, not just read.

Log in to unlockSee a live demo →

Fig. 1Interactive explainer · premium

What it is

DomainShuttle is a subject-driven text-to-video system built on the Wan2.1/2.2 DiT video models that adds three components: Domain-MoT (separate transformer branches for reference images versus video plus a domain-aware AdaLN modulation), Video-Reference DualRoPE (placing reference tokens in a separate positional encoding space from video tokens), and a Cross-Pair Consistent Loss that aligns two different reference sets of the same subject. The goal is to preserve a reference subject's intrinsic features while still allowing style, domain, and semantic changes driven by the text prompt (cross-domain generation), rather than only copying the reference (in-domain).

Why it matters

For anyone building video personalization tools (advertising, creative design, AI filmmaking), existing methods tend to copy-paste the reference subject and resist stylistic transformation, so you cannot easily put a real person into a watercolor or 3D-animation scene while keeping their identity. This method aims to keep identity fidelity and still follow style instructions, which is the practical difference between a rigid template filler and a flexible editor.

Practical takeaway

Watch for this in Wan-based video generation tooling: it is fine-tuned on top of open-source Wan2.1/2.2-14B, so if you already use those models, the approach (decoupled reference branch, separate RoPE space for reference tokens, consistency loss across reference sets) is a recipe you could adopt for subject-driven generation that needs style transfer, not just identity locking.

Key result

On the authors' own 220-sample test set (110 in-domain, 110 cross-domain), the Wan2.2 version reports a Cross-Domain Score of 0.861 versus 0.725 for the best baseline (Kling 1.6), an 18.7% relative improvement. Note the evaluation set is self-constructed and modest in size, and several cross-domain metrics rely on the authors' own MLLM-based scoring (GPT-5.2, Qwen3-VL) and image-edit-based CLIP pipelines rather than an established public benchmark.

Subscribe

Get the next issue.

Free. One email a week. Unsubscribe any time: no account, no dark patterns.