Fig.1

From Issue #6 · 2026-06-15

ABot-Earth 0.5: Generative 3D Earth Model

Ming Qian, Tianjian Ouyang, Mingchao Sun, Zijian Wang, Jincheng Xiong, Jiarong Han, Yongchang Zhang, Jiawei Zhang, Xu Wang, Yu Liu, Luyang Tang, Fei Yu, Zengye Ge, Mengmeng Du, Yuan Liu, Nianfei Fan, Song Wang, Yingliang Peng, Chunxue Jia, Yang Liu, Shiying Zeng, Haozhe Shi, Junnan Lai, Hongyu Pan, Zheng Wu, Ning Guo, Mu Xu, Hang Zhang

arXiv:2606.09967 · 486▲ · cs.CV

View on arXiv →

Premium readers get an interactive explainer for this paper — a figure you can poke at, not just read.

Log in to unlockSee a live demo →

Fig. 1Interactive explainer · premium

What it is

ABot-Earth 0.5 is a generative 3D system that synthesizes city-scale outdoor environments directly in the 3D Gaussian Splatting (3DGS) representation, conditioned on ordinary satellite imagery. It trains a compression-generation model on real-world 3DGS reconstructions (built by their own ABot-3DGS pipeline) and generates tiled scenes with a sliding-window inference scheme and native multi-level-of-detail output for streaming.

Why it matters

Traditional planetary 3D reconstruction relies on oblique photogrammetry and LiDAR, which is expensive and takes months to years to update. A generative approach that produces plausible 3D from satellite images at roughly 10 minutes per square kilometer changes the economics of covering unmapped regions, though the output is synthesized (plausible) rather than surveyed, so it should not be treated as ground truth.

Practical takeaway

Watch for satellite-to-3DGS generation as a way to fill 3D coverage gaps where photogrammetry does not exist (they show a synthesized Ireland where Google Earth falls back to flat 2D). The output is native 3DGS on open 3D Tiles standards, so it can be composited with separately reconstructed high-fidelity landmarks and used as a simulation sandbox for UAV navigation.

Key result

Reported FID of 16.1 versus 69.5 for EarthCrafter (the prior best baseline) on 2D renderings, with KID 0.006 vs 0.061. Major caveat stated by the authors themselves: baseline FID/KID were computed against different ground-truth sets and different camera poses, so the numbers are 'for reference only' rather than a controlled comparison.

Subscribe

Get the next issue.

Free. One email a week. Unsubscribe any time: no account, no dark patterns.