Vision
- ABot-Earth 0.5: Generative 3D Earth Model
Ming Qian, Tianjian Ouyang, Mingchao Sun, Zijian Wang, Jincheng Xiong, Jiarong Han, Yongchang Zhang, Jiawei Zhang, Xu Wang, Yu Liu, Luyang Tang, Fei Yu, Zengye Ge, Mengmeng Du, Yuan Liu, Nianfei Fan, Song Wang, Yingliang Peng, Chunxue Jia, Yang Liu, Shiying Zeng, Haozhe Shi, Junnan Lai, Hongyu Pan, Zheng Wu, Ning Guo, Mu Xu, Hang Zhang
ABot-Earth 0.5 is a generative 3D system that synthesizes city-scale outdoor environments directly in the 3D Gaussian Splatting (3DGS) representation, conditioned on ordinary satellite imagery. It trains a compression-generation model on real-world 3DGS reconstructions (built by their own ABot-3DGS pipeline) and generates tiled scenes with a sliding-window inference scheme and native multi-level-of-detail output for streaming.
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, Neil Houlsby
This paper introduces the Vision Transformer (ViT), which applies a standard Transformer encoder directly to image classification by splitting an image into fixed-size patches (for example 16x16 pixels), linearly embedding each patch as a token, adding position embeddings, and processing the sequence with self-attention. It deliberately avoids convolutions and image-specific inductive biases except for the initial patch extraction and position embedding interpolation at fine-tuning time.
- Auto-Encoding Variational Bayes
Diederik P Kingma, Max Welling
This paper introduces the reparameterization trick (rewriting a latent variable z as a deterministic function of the parameters plus fixed noise) so that the variational lower bound becomes differentiable and trainable with ordinary stochastic gradient descent. Applying this to an encoder/decoder pair of neural networks gives the variational auto-encoder (VAE), trained end to end with the AEVB algorithm.
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe, Christian Szegedy
This paper introduces Batch Normalization, a layer inserted before each nonlinearity that normalizes activations to zero mean and unit variance using per-mini-batch statistics, then applies a learned scale and shift. The normalization is part of the network architecture so gradients backpropagate through it, and at inference time it uses population statistics computed from moving averages.
- Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generation
Min Zhao, Hongzhou Zhu, Kaiwen Zheng, Zihan Zhou, Bokai Yan, Xinyuan Li, Xiao Yang, Chongxuan Li, Jun Zhu
The paper distills bidirectional video diffusion models into few-step autoregressive students for real-time interactive video generation, replacing the expensive causal ODE initialization step of prior work (Causal Forcing) with causal consistency distillation (causal CD). Causal CD gets its training signal from a single online teacher ODE step between adjacent timesteps on real videos, avoiding the need to precompute and store full PF-ODE trajectories.
- CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
Dongsheng Ma, Jiayu Li, Zhengren Wang, Yijie Wang, Jiahao Kong, Weijun Zeng, Jutao Xiao, Jie Yang, Wentao Zhang, Bin Wang, Conghui He
CiteVQA is a document VQA benchmark that requires models to return element-level bounding-box citations alongside each answer, then scores both jointly with a metric called Strict Attributed Accuracy (SAA), which only credits a sample when the answer is correct AND the cited region matches ground truth. Ground-truth citations are generated by an automated pipeline (document parsing plus MLLM agents plus masking-based ablation to find crucial evidence) and validated by expert review across 1,897 questions from 711 multi-page PDFs.
- Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, Alisson Azzolini, Junjie Bai, Maciej Bala, Yogesh Balaji, Josh Bapst, Aarti Basant, Mukesh Beladiya, Mohammad Qazim Bhat, Zaid Pervaiz Bhat, Dan Blick, Vanni Brighella, Han Cai, Tiffany Cai, Eric Cameracci, Jiaxin Cao, Yulong Cao, Mark Carlson, Carlos Casanova, Ting-Yun Chang, Yan Chang, Yu-Wei Chao, Prithvijit Chattopadhyay, Roshan Chaudhari, Chieh-Yun Chen, Junyu Chen, Ke Chen, Qizhi Chen, Wenkai Chen, Xiaotong Chen, Yu Chen, An-Chieh Cheng, Click Cheng, Xiu Chia, Jeana Choi, Chaeyeon Chung, Wenyan Cong, Yin Cui, Magdalena Dadela, Nalin Dadhich, Wenliang Dai, Joyjit Daw, Alperen Degirmenci, Rodrigo Vieira Del Monte, Robert Denomme, Sameer Dharur, Marco Di Lucca, Ke Ding, Wenhao Ding, Yifan Ding, Yuzhu Dong, Nicole Drumheller, Yilun Du, Aigul Dzhumamuratova, Aleksandr Efitorov, Hamid Eghbalzadeh, Naomi Eigbe, Imad El Hanafi, Hassan Eslami, Benedikt Falk, Jiaojiao Fan, Jim Fan, Amol Fasale, Sergiy Fefilatyev, Liang Feng, Francesco Ferroni, Sanja Fidler, Xiao Fu, Vikram Fugro, Prashant Gaikwad, TJ Galda, Katelyn Gao, Yihuai Gao, Wenhang Ge, Sreyan Ghosh, Arushi Goel, Vivek Goel, Akash Gokul, Rama Govindaraju, Jinwei Gu, Miguel Guerrero, Elfie Guo, Aryaman Gupta, Siddharth Gururani, Hugo Hadfield, Song Han, Ankur Handa, Zekun Hao, Mohammad Harrim, Ali Hassani, Nathan Hayes-Roth, Yufan He, Chris Helvig, Cyrus Hogg, Madison Huang, Michael Huang, Sophia Huang, Yufan Huang, Jacob Huffman, DeLesley Hutchins, Suneel Indupuru, Boris Ivanovic, Arihant Jain, Joel Jang, Ryan Ji, Yanan Jian, Dongfu Jiang, Jingyi Jin, Atharva Joshi, Nikhilesh Joshi, Pranjali Joshi, Andy Ju, Jaehun Jung, Weiwei Kang, Scott Kassekert, Jan Kautz, Ashna Khetan, Julia Kiczka, Slawek Kierat, Gwanghyun Kim, Kuno Kim, Sunny Kim, Kezhi Kong, Xin Kong, Zhifeng Kong, Tomasz Kornuta, Egor Krivov, Hui Kuang, Saurav Kumar, Chia-Wen Kuo, George Kurian, Wojciech Kutak, JF Lafleche, Himangshu Lahkar, Omar Laymoun, Jayjun Lee, Sanggil Lee, Gabriele Leone, Boyi Li, Freya Li, Jiajun Li, Jinfeng Li, Ling Li, Pengcheng Li, Shangru Li, Tingle Li, Xiaolong Li, Xuan Li, Zhaoshuo Li, Zhiqi Li, Hao Liang, Maosheng Liao, Chen-Hsuan Lin, Tsung-Yi Lin, Ming-Yu Liu, Sifei Liu, Zihan Liu, Hai Loc Lu, Xiangyu Lu, Alice Luo, Ruipu Luo, Wenjie Luo, Jiangran Lyu, Martin Ding Ma, Nic Ma, Qianli Ma, Dawid Majchrowski, Louis Marcoux, Miguel Martin, Qing Miao, Ashkan Mirzaei, Shreyas Misra, Kaichun Mo, Durra Mohsin, Hyejin Moon, Pawel Morkisz, Saeid Motiian, Kirill Motkov, Seungjun Nah, Yashraj Narang, Deepak Narayanan, Thabang Ngazimbi, Julian Ouyang, Shubham Pachori, David Page, Yatian Pang, Sehwi Park, Mahesh Patekar, Mostofa Patwary, Marco Pavone, Trung Pham, Wei Ping, Soha Pouya, Shrimai Prabhumoye, Varun Praveen, Delin Qu, Hesam Rabeti, Morteza Ramezanali, Marilyn Reeb, Xuanchi Ren, Kristen Rumley, Wojciech Rymer, Jun Saito, Yeongho Seol, John Shao, Piyush Shekdar, Tianwei Shen, Humphrey Shi, Min Shi, Stella Shi, Kevin Shih, Mohammad Shoeybi, Mateusz Sieniawski, Shuran Song, Alexander Sotelo, Amir Sotoodeh, Sunil Srinivasa, Vignesh Srinivasakumar, Bartosz Stefaniak, Rahul Heinrich Steiger, Shangkun Sun, Jiaxiang Tang, Shitao Tang, Yangyang Tang, Yue Tang, Tolou Tavakkoli, Kayley Ting, Krzysztof Tomala, Wei-Cheng Tseng, Jibin Varghese, Sergei Vasilev, Thomas Volk, Raju Wagwani, Roger Waleffe, Andrew Z. Wang, Boxiang Wang, Haoxiang Wang, Qiao Wang, Shihao Wang, Shijie Wang, Ting-Chun Wang, Yan Wang, Yu Wang, Rohit Watve, David Wehr, Fangyin Wei, Xinshuo Weng, Jay Zhangjie Wu, Kedi Wu, Hongchi Xia, Summer Xiao, Tianjun Xiao, Kevin Xie, Daguang Xu, Jiashu Xu, Mengyao Xu, Ruqing Xu, Xingqian Xu, Yao Xu, Dinghao Yang, Dong Yang, Hans Yang, Xiaodong Yang, Xuning Yang, Yichu Yang, Yurong You, Zhiding Yu, Hao Yuan, Simon Yuen, Xiaohui Zeng, Pengcuo Zeren, Cindy Zha, Haotian Zhang, Jenny Zhang, Jing Zhang, Liangkai Zhang, Paris Zhang, Shun Zhang, Xuanmeng Zhang, Zhizheng Zhang, Ann Zhao, Yilin Zhao, Yuliya Zhautouskaya, Charles Zhou, Fengzhe Zhou, Shilin Zhu, Yuke Zhu, Dima Zhylko, Artur Zolkowski
Cosmos 3 is a family of omnimodal models (4B, 16B, 64B params) that jointly process and generate language, image, video, audio, and action within a single Mixture-of-Transformers architecture. It uses a dual-tower design where one parameter set handles autoregressive reasoning (next-token prediction) and a separate set handles diffusion-based generation of pixels/audio/actions, with the two towers sharing a joint attention operation and both initialized from a pretrained Qwen3-VL.
- DanceOPD: On-Policy Generative Field Distillation
Wei Zhou, Xiongwei Zhu, Zelin Xu, Bo Dong, Lixue Gong, Yongyuan Liang, Meng Chu, Leigang Qu, Lingdong Kong, Wei Liu, Tat-Seng Chua
DanceOPD is an on-policy distillation method for flow-matching image generators that combines multiple frozen capability models (text-to-image, local editing, global editing) into one student. Each training sample is hard-routed to exactly one teacher's velocity field, that field is queried on a single low-noise state from the student's own rollout (stop-gradient), and the student is trained with plain velocity MSE.
- Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, Pieter Abbeel
This paper introduces Denoising Diffusion Probabilistic Models (DDPMs), a generative model that learns to reverse a fixed Markov chain that gradually adds Gaussian noise to images. The key practical simplification is training a U-Net to predict the noise added at each timestep using a plain weighted mean-squared-error objective, which the authors show is equivalent to denoising score matching over multiple noise levels.
- DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation
Nan Chen, Yiyang Cai, Rongchang Xie, Junwen Pan, Cheng Chen, Weinan Jia, Zhuowei Chen, Wen Zhou, Zhenbang Sun, Wenhan Luo
DomainShuttle is a subject-driven text-to-video system built on the Wan2.1/2.2 DiT video models that adds three components: Domain-MoT (separate transformer branches for reference images versus video plus a domain-aware AdaLN modulation), Video-Reference DualRoPE (placing reference tokens in a separate positional encoding space from video tokens), and a Cross-Pair Consistent Loss that aligns two different reference sets of the same subject. The goal is to preserve a reference subject's intrinsic features while still allowing style, domain, and semantic changes driven by the text prompt (cross-domain generation), rather than only copying the reference (in-domain).
- DreamX-World 1.0: A General-Purpose Interactive World Model
DreamX Team, Yancheng Bai, Rui Chen, Xiangxiang Chu, Rujing Dang, Hao Dou, Bingjie Gao, Qiwen Gu, Siyu Hong, Jiachen Lei, Geng Li, Jifan Li, Ruimin Lin, Qingfeng Shi, Bingze Song, Lei Sun, Jing Tang, Ruitian Tian, Jun Wang, Jiahong Wu, Pengfei Zhang, Shen Zhang, Jiashu Zhu
DreamX-World 1.0 is an interactive text/image-to-video world model built by fine-tuning Wan2.2 to support camera navigation, revisiting earlier scenes, and prompted multi-object events across photorealistic, game, and stylized domains. It combines a camera-conditioning method called E-PRoPE (projective positional encoding applied to spatially downsampled tokens), geometry-based memory retrieval, DMD distillation into a few-step autoregressive generator, and RL alignment, with serving optimizations to hit real-time streaming.
- Flow-OPD: On-Policy Distillation for Flow Matching Models
Zhen Fang, Wenxuan Huang, Yu Zeng, Yiming Zhao, Shuang Chen, Kaituo Feng, Yunlong Lin, Lin Chen, Zehui Chen, Shaosheng Cao, Feng Zhao
Flow-OPD is a post-training method for text-to-image flow matching models that replaces multi-reward reinforcement learning with on-policy distillation from multiple single-task teacher models. It trains domain-expert teachers via single-reward GRPO, then distills them into one student by routing prompts to the relevant teacher and matching velocity fields, showing that the reverse-KL divergence between student and teacher policies reduces analytically to a weighted MSE loss on the vector fields.
- Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
Fangfu Liu, Kai He, Tianchang Shen, Tianshi Cao, Sanja Fidler, Yueqi Duan, Jun Gao, Igor Gilitschenski, Zian Wang, Xuanchi Ren
Gamma-World is a video world model that generates synchronized, action-controllable video streams for multiple agents (game players or robot arms) sharing one environment. It introduces two components: Simplex Rotary Agent Encoding, a parameter-free extension of 3D RoPE that assigns each agent a rotary phase at the vertex of a regular simplex so agents stay distinct but permutation-symmetric, and Sparse Hub Attention, where learnable hub tokens mediate cross-agent communication instead of dense all-to-all attention.
- Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio
This is the original Generative Adversarial Nets (GAN) paper. It trains two neural networks simultaneously in a minimax game: a generator G that maps random noise to fake samples, and a discriminator D that tries to tell real training data from generated data, with both trained by ordinary backpropagation.
- Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, Ruslan R. Salakhutdinov
This is the original dropout paper: during training, each hidden unit is randomly omitted with probability 0.5 on every training case, which stops units from co-adapting to rely on specific other units. At test time you use the full network with outgoing weights halved, which approximates averaging over the exponentially many thinned networks.
- Learning Transferable Visual Models From Natural Language Supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, Ilya Sutskever
This is the CLIP paper. It trains an image encoder and a text encoder jointly on 400 million (image, text) pairs scraped from the internet, using a contrastive objective that predicts which caption goes with which image within a batch rather than predicting exact caption words. After pre-training, you build a classifier for any dataset by feeding the class names (as text prompts) through the text encoder, so the model classifies images it was never explicitly trained to label.
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
Shihao Wang, Shilong Liu, Yuanguo Kuang, Xinyu Wei, Yangzhou Liu, Zhiqi Li, Yunze Man, Guo Chen, Andrew Tao, Guilin Liu, Jan Kautz, Lei Zhang, Zhiding Yu
LocateAnything is a vision-language model for object detection and grounding that uses Parallel Box Decoding (PBD), which predicts all four coordinates of a bounding box in a single forward step instead of generating them token by token. It treats each box (or point) as an atomic block with bidirectional intra-block attention, trained jointly with a standard next-token prediction objective and offering three inference modes (Fast, Slow, Hybrid).
- Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers
Pengqi Lu
The paper diagnoses a training collapse in very deep Diffusion Transformers where token representations homogenize toward their sequence mean, which the author calls Mean Mode Screaming, and traces it to a gradient decomposition where a mean-coherent component grows as O(T) once tokens align. It proposes MV-Split Residuals, a modified Post-Norm residual merge that applies separate learnable gains to the centered (token-varying) part of the residual branch versus a leaky replacement of the trunk mean.
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
Kangsheng Duan, Ziyang Xu, Wenyu Liu, Xiaohu Ruan, Xiaoxin Chen, Xinggang Wang
Moebius is a 0.22B-parameter image inpainting model built on a latent diffusion U-Net whose transformer blocks are replaced with a Local-Lambda Mix Interaction (LLambdaMI) block that summarizes spatial context and semantic priors into fixed-size linear matrices instead of quadratic attention. It is trained with an adaptive multi-granularity knowledge distillation scheme that aligns the small student to the larger PixelHacker teacher entirely in latent space using gradient-norm-balanced losses.
- OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
Jiwen Liu, Shujuan Li, Zhixue Fang, Xiaohan Li, Yan Zhou, Zijie Meng, Zhimin Zhang, Yawen Luo, Guoxin Zhang, Yu-Shen Liu, Pengfei Wan
OmniDirector is a video generation framework that clones camera motion from a reference video by rendering the extracted camera poses as a 'camera grid' (a video of grid lines moving through an empty 3D room), then feeding that grid into a Multi-Modal Diffusion Transformer via token concatenation. It also uses a hierarchical prompt expansion agent built on Qwen3-VL to describe camera motion (split into inter-shot and intra-shot descriptions) and fuse it with the reference image and user prompt at inference time.
- Orca: The World is in Your Mind
Yihao Wang, Yuheng Ji, Mingyu Cao, Yanqing Shen, Runze Xiao, Huaihai Lyu, Senwei Xie, Euan Liu, Klara Tian, Tianfeng Long, Yichi Zhang, Zhengliang Cai, Ruike Chen, Jifan Zhao, Ruochuan Shi, Zihan Tang, Jing Lyu, Wenxing Tan, Ningbo Zhang, Yangtao Hu, Yuming Gao, Xiansheng Chen, Junkai Zhao, Congsheng Xu, Boan Zhu, Ziqi Wang, Yupu Feng, Qiongqiong Zhang, Yingli Zhao, Yulong Ao, Shaoxuan Xie, You Liu, Guocai Yao, Leiduo Zhang, Xiaodan Liu, Yunyan Zhang, Yance Jiao, Xinyan Yang, Jiaxing Wei, Xu Liu, Tengfei Pan, Shaokai Nie, Chunlei Men, Sen Cui, Xiaojie Jin, Hongyang Li, Jianlan Luo, Yao Mu, Yunchao Wei, Jun Yan, Hang Zhao, Xiaolong Zheng, Jiaming Li, Yonghua Lin, Tiejun Huang, Zhongyuan Wang, Pengwei Wang
Orca is an encoder-decoder model that learns a shared 'world latent space' from video and language by predicting the latent representation of adjacent frames (unconscious learning) and event-conditioned future states described by text (conscious learning), plus a standard VQA loss. After pretraining, the backbone (built on a Qwen VLM) is frozen and small task-specific decoders are trained for text, image prediction, and robot action generation to test whether the shared latent transfers.
- PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models
Yueyi Sun, Yuhao Wang, Jason Li, Ye Tian, Tao Zhang, Jacky Mai, Yihan Wang, Haochen Wang, Jinbin Bai, Ling Yang, Yunhai Tong
PerceptionDLM is a multimodal diffusion language model that captions multiple image regions at once inside a single denoising pass, instead of the region-by-region autoregressive decoding used by prior models. It combines a SigLIP-2 vision encoder and an 8B LLaDA diffusion backbone with region prompting (learnable per-region embeddings), RoI-aligned feature replay, and a structured attention mask that isolates each region's tokens while sharing global context.
- Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
Caixin Kang, Tianyu Yan, Sitong Gong, Mingfang Zhang, Liangyang Ouyang, Ruicong Liu, Bo Zheng, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang
The paper introduces Grounded Personality Reasoning (GPR), a task that forces multimodal LLMs to justify each Big Five (OCEAN) personality rating with cited, timestamped behavioral evidence, and releases MM-OCEAN, a benchmark of 1,104 videos and 5,320 cue-grounding multiple-choice questions built by a multi-agent annotation pipeline with human verification. It benchmarks 27 MLLMs across three tiers (rating, open-ended reasoning, structured cue grounding) plus four failure-mode metrics that separate 'right answer' from 'right reasons'.
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Haiwen Diao, Penghao Wu, Hanming Deng, Jiahao Wang, Shihao Bai, Silei Wu, Weichen Fan, Wenjie Ye, Wenwen Tong, Xiangyu Fan, Yan Li, Yubo Wang, Zhijie Cao, Zhiqian Lin, Zhitao Yang, Zhongang Cai, Yuwei Niu, Yue Zhu, Bo Liu, Chengguang Lv, Haojia Yu, Haozhe Xie, Hongli Wang, Jianan Fan, Jiaqi Li, Jiefan Lu, Jingcheng Ni, Junxiang Xu, Kaihuan Liang, Lianqiang Shi, Linjun Dai, Linyan Wang, Oscar Qian, Peng Gao, Pengfei Liu, Qingping Sun, Rui Shen, Ruisi Wang, Shengnan Ma, Shuang Yang, Siyi Xie, Siying Li, Tianbo Zhong, Xiangli Kong, Xuanke Shi, Yang Gao, Yongqiang Yao, Yves Wang, Zhengqi Bai, Zhengyu Lin, Zixin Yin, Wenxiu Sun, Ruihao Gong, Quan Wang, Lewei Lu, Lei Yang, Ziwei Liu, Dahua Lin
SenseNova-U1 is a multimodal model that handles both understanding (image/text comprehension) and image generation in a single architecture that operates directly on raw pixels and text, dropping the usual pretrained vision encoder and VAE. It uses a Mixture-of-Transformers (MoT) backbone where a shared attention path processes both streams, training understanding with next-token prediction and generation with pixel-space flow matching. Two variants ship: a dense 8B and a 30B mixture-of-experts model with about 3B active parameters.
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Seokju Cho, Ryo Hachiuma, Abhishek Badki, Hang Su, Byung-Kwan Lee, Chan Hee Song, Sifei Liu, Subhashree Radhakrishnan, Seungryong Kim, Yu-Chiang Frank Wang, Min-Hung Chen
SpatialClaw is a training-free agent framework that gives a vision-language model a persistent Python kernel (preloaded with input frames, perception tools like SAM3 and Depth Anything 3, and libraries like NumPy/SciPy) and lets it write and run one code cell per step, inspecting intermediate results before writing the next cell. This 'code as the action interface' approach replaces two prior designs: single-pass code (commit to a full program before seeing any output) and structured JSON/XML tool-calls (fixed API menu).
- Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
Bin Wu, Mengqi Huang, Shaojin Wu, Weinan Jia, Yuxin Wang, Zhendong Mao, Yongdong Zhang
Stream-R1 modifies distribution matching distillation (DMD) for streaming autoregressive video generation by reweighting the distillation loss with a pretrained video reward model at two levels: a per-rollout scalar weight (exponential of the reward score) and a per-pixel spatiotemporal weight derived from backpropagating the reward model's gradient. It replaces the uniform weighting that treats every rollout, frame, and pixel as equally reliable supervision with reward-guided weighting that concentrates optimization on regions where refinement helps most.
- Stream-T1: Test-Time Scaling for Streaming Video Generation
Yijing Tu, Shaojin Wu, Mengqi Huang, Wenchuan Wang, Yuxin Wang, Chunxiao Liu, Zhendong Mao
Stream-T1 is a test-time scaling framework built on top of a streaming (chunk-by-chunk autoregressive) video diffusion model (LongLive). It combines three inference-time tricks: initializing each chunk's noise by spherical interpolation from the previous chunk's noise, beam search pruning of chunk candidates using a combined image-reward plus video-reward score, and reward-guided routing of evicted KV cache entries into discard, EMA-merge, or append pathways.
- WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
Kaining Ying, Hengrui Hu, Siyu Ren, Jiamu Li, Fengjiao Chen, Ziwen Wang, Xuezhi Cao, Xunliang Cai, Henghui Ding
WBench is a benchmark for evaluating interactive video world models (systems that generate the next video frame conditioned on observation history plus a user action). It defines 289 test cases with 1,058 multi-turn interaction turns across five dimensions (video quality, setting adherence, interaction adherence, consistency, physics compliance), scored by 22 automatic sub-metrics that combine specialist vision models (MegaSaM pose estimation, SAM2, Depth Anything 3, DINOv2) with VLM-based grading.