Data
- Are We Ready For An Agent-Native Memory System?
Wei Zhou, Xuanhe Zhou, Shaokun Han, Hongming Xu, Guoliang Li, Zhiyu Li, Feiyu Xiong, Fan Wu
A benchmark study that decomposes 12 LLM agent memory systems (MemGPT/Letta, Mem0, Zep, A-MEM, MemOS, MemoryOS, Cognee, etc.) into four modules (representation/storage, extraction, retrieval/routing, and maintenance) and evaluates them from a data management angle. It runs end-to-end and ablation experiments across five workloads spanning 11 datasets, measuring not just task accuracy but retrieval fidelity, update robustness, long-horizon stability, and operational cost (index build time, query latency).
- Beyond IID: How General Are Tabular Foundation Models, Really?
Lennart Purucker, Andrej Tschalzev, Nick Erickson, Gioia Blayer, David Holzmüller, Alan Arazi, Alexander Pfefferle, Mustafa Tajjar, Gaël Varoquaux, Frank Hutter
The paper introduces BeyondArena, a benchmark of 142 manually curated tabular datasets that spans IID, temporal, and grouped (non-IID) prediction tasks across sample sizes from 100 to 1 million rows, plus DataFoundry, a Python framework for reproducible dataset curation. It evaluates 11 models including three tabular foundation models (TabPFN-2.6, TabICLv2, TabDPT) against gradient-boosted trees and MLPs using in-context learning for the foundation models and tuning plus ensembling for the traditional ones.
- EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions
Jincheng Zhong, Weizhi Wang, Che Jiang, Kai Tian, Zhenzhao Yuan, Junlin Yang, Dianqiao Lei, Kaiyan Zhang
EnterpriseClawBench is a benchmark built by an automated pipeline that converts real internal workplace agent sessions into 852 reproducible tasks, each with recovered input fixtures, rewritten single-turn prompts, role/skill taxonomy labels, hard delivery rules, and semantic rubrics. It evaluates harness-model combinations (Claude Code, Codex, DeepAgents, Hermes, OpenClaw paired with various models) on producing usable business artifacts, scoring both objective delivery rules and LLM-judged quality across five dimensions.
- K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
Nahyun Lee, Dongkeun Yoon, Guijin Son, Geewook Kim, Dayoon Ko, Jeonghun Park, Haneul Yoo, Jaewon Cho, Junghun Park, Changyoon Lee, Kyochul Jang, Jaeyeon Kim, Eunsu Kim, Woojin Cho, Seungone Kim
K-BrowseComp is a web-browsing agent benchmark of 400 problems grounded in Korean web contexts: a 300-question human-verified subset built by native Korean annotators, plus a 100-question synthetic split generated by a browsing agent (Claude Code) that constructs questions backwards from seed pages and filters them against a taxonomy of failure modes. Each question has a single stable answer requiring multi-hop or parallel-constraint retrieval across multiple Korean websites, and models are evaluated pass@1 under a fixed 10-call search budget.
- Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
Zhifei Xie, Kaiyu Pang, Haobin Zhang, Deheng Ye, Xiaobin Hu, Shuicheng Yan, Chunyan Miao
Mega-ASR fine-tunes Qwen3-ASR-1.7B to handle heavily degraded real-world audio using two techniques: Acoustic-to-Semantic Progressive Supervised Fine-Tuning (a WER-graded curriculum that trains the encoder, then the LLM, then jointly) and Dual-Granularity WER-Gated Policy Optimization (a reinforcement learning reward that switches between token-level and sentence-level scoring based on WER). It is trained on Voices-in-the-Wild-2M, a 2.4M-clip synthetic dataset built by simulating 7 atomic acoustic effects and 54 compound scenarios at the spectrogram level.
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
Alan Arazi, Eilam Shapira, Shoham Grunblat, Mor Ventura, Elad Hoffer, Gioia Blayer, David Holzmüller, Lennart Purucker, Gaël Varoquaux, Frank Hutter, Roi Reichart
MulTaBench is a benchmark of 40 datasets (20 image-tabular, 20 text-tabular) for multimodal tabular learning, curated by an automated pipeline that only keeps datasets where each modality adds independent predictive signal and where task-specific tuning of the encoder beats frozen off-the-shelf embeddings. The curation compares four conditions (unimodal, joint frozen, joint target-aware) across five tabular learners, using LoRA finetuning of the top 3 layers of e5 (text) and DINO-v3 (image) encoders as a preprocessing step to produce Target-Aware Representations.
- NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
Yuru Wang, Lejun Cheng, Yuxin Zuo, Sihang Zeng, Bingxiang He, Che Jiang, Junlin Yang, Yuchong Wang, Kaikai Zhao, Weifeng Huang, Kai Tian, Zhenzhao Yuan, Jincheng Zhong, Weizhi Wang, Ning Ding, Bowen Zhou, Kaiyan Zhang
NatureBench is a benchmark of 90 tasks distilled from Nature-family science papers (2022-2025), built by an automated pipeline (NatureGym) that turns each paper into a containerized package with a task brief, data, a hidden test set, and an automated evaluator, while removing the original method so agents must find their own solution. It scores coding agents against each paper's published state of the art using a SOTA-normalized relative gap metric, with a post-hoc judge that flags shortcuts like output fabrication and feedback gaming.
- OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories
Yuwen Du, Rui Ye, Shuo Tang, Keduan Huang, Xinyu Zhu, Yuzhu Cai, Siheng Chen
The paper trains a 30B search agent (OpenSeeker-v2, based on Qwen3-30B-A3B-Thinking) using only supervised fine-tuning on synthetic ReAct trajectories, skipping the usual continual pre-training and reinforcement learning stages. The main contribution is three tweaks to the data synthesis pipeline: expanding the knowledge graph size used to generate questions, enlarging the tool set, and filtering out any trajectory solvable in too few tool-call steps to enforce a difficulty floor.
- PhysBrain 1.0 Technical Report
Shijie Lian, Bin Yu, Xiaopeng Lin, Changti Wu, Hang Yuan, Xiaolin Hu, Zhaolong Shen, Yuzhuo Miao, Haishan Liu, Yuxuan Tian, Yukun Shi, Cong Huang, Kai Chen
PhysBrain 1.0 is a vision-language-action (VLA) pipeline that converts large-scale egocentric human video into structured physical supervision. A data engine parses clips into JSON scene records (objects, spatial dynamics, action execution, depth relations), renders those into natural-language question-answer pairs to train a base VLM, then adapts that model to robot control with a design meant to preserve general multimodal ability.
- TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
Hanyu Guo, Jiedong Yang, Chao Chen, Longfei Xu, Kaikui Liu, Xiangxiang Chu
TransitLM is a dataset of over 13 million transit route planning records (from Amap logs across four Chinese cities, 120,845 stations) plus a three-task benchmark, released as a continual pre-training corpus and supervised fine-tuning data. The authors validate it by continually pre-training then fine-tuning Qwen3 models (0.6B to 4B) that generate complete transit routes end-to-end from origin-destination GPS coordinates, with each station ID registered as a dedicated vocabulary token so routes cannot be hallucinated via character composition.