Fig.1

From Issue #8 · 2026-06-29

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Lianghua Huang, Zhi-Fan Wu, Wei Wang, Yupeng Shi, Mengyang Feng, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuzheng Wang, Zoubin Bi

arXiv:2606.25041 · 115▲ · cs.CV, cs.AI, cs.GR, cs.SD

View on arXiv →

Premium readers get an interactive explainer for this paper — a figure you can poke at, not just read.

Log in to unlockSee a live demo →

Fig. 1Interactive explainer · premium

What it is

Wan-Streamer is a single Transformer that handles text, audio, and video as both input and output, using block-causal attention plus conditional flow matching to generate synchronized speech and video responses in a streaming, full-duplex manner. It replaces the usual cascade of separate VAD, ASR, LLM, TTS, and avatar-rendering modules with one end-to-end model built around causal VAEs, causal encoders/decoders, and a thinker-performer split at inference time.

Why it matters

Cascaded voice/avatar systems add waiting time at every module boundary and accumulate recognition and synchronization errors. A unified causal model lets perception, response timing, interruption handling, and audio-visual sync be learned together, which the authors argue is the only way to reach genuinely low-latency full-duplex behavior with a live video agent rather than a speech-only or text-first pipeline.

Practical takeaway

Watch for interactive agents that both perceive your video and generate a synchronized talking avatar from one model instead of stitched components. Note this is a v0.1 proof of concept: outputs are only 192p and there are no accuracy or quality benchmarks against competing systems, just latency and runtime scope comparisons.

Key result

About 200 ms model-side response latency and about 550 ms total interaction latency (including a 350 ms bidirectional network budget) with 25 FPS video output, on a two-GPU thinker-performer serving path. The comparison tables are heavily caveated by the authors themselves: rival systems report different endpoints (first-packet, first-token, endpointing, API TTFB), so the numbers are not directly comparable, and no output-quality metrics are provided.

Subscribe

Get the next issue.

Free. One email a week. Unsubscribe any time: no account, no dark patterns.