WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
Kaining Ying, Hengrui Hu, Siyu Ren, Jiamu Li, Fengjiao Chen, Ziwen Wang, Xuezhi Cao, Xunliang Cai, Henghui Ding
arXiv:2605.25874 · 104▲ · cs.CV
View on arXiv →Premium readers get an interactive explainer for this paper — a figure you can poke at, not just read.
What it is
WBench is a benchmark for evaluating interactive video world models (systems that generate the next video frame conditioned on observation history plus a user action). It defines 289 test cases with 1,058 multi-turn interaction turns across five dimensions (video quality, setting adherence, interaction adherence, consistency, physics compliance), scored by 22 automatic sub-metrics that combine specialist vision models (MegaSaM pose estimation, SAM2, Depth Anything 3, DINOv2) with VLM-based grading.
Why it matters
World model evaluation has been fragmented across task-specific protocols, making fair comparison impossible. WBench unifies navigation control across text, 6-DoF pose, and discrete keyboard action so models with different native input interfaces can be scored on a shared subset, and it separates 'what the world is' from 'what the user requests' so failure modes (e.g. correct rendering but ignored actions, or drift across turns) can be localized. This is primarily a tool for researchers and model developers rather than something that changes a production workflow.
Practical takeaway
If you are choosing or building an interactive world model, expect no single system to win across all five dimensions: navigation controllability is decoupled from rendering quality, and high consistency scores are often inflated by scenes that barely move (use the gated spatial consistency metric to catch this). Open-source models (HY-World 1.5, LingBot-World, Matrix-Game 3.0) are competitive with closed-source ones on specific axes.
Key result
Across 20 models on the 158-case navigation split, dedicated camera-controlled (76.0) and action-conditioned (77.7) models beat text-driven models (67.6) on navigation by about 10 points, yet text-driven models led on physics compliance (67.0 vs 64.2 and 61.7). Navigation degraded roughly 33 points from turn 1 to turn 4+, the fastest decay of any interaction type. Automated scores matched human blind pairwise rankings with Spearman rho >= 0.94 across all ten aspects, though this validates ranking order, not absolute accuracy.
Subscribe
Get the next issue.
Free. One email a week. Unsubscribe any time: no account, no dark patterns.