Depth Anything 3: Recovering the Visual Space from Any Views

ICLROral2026

Authors
Haotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen, Zhenyu Li, Yang Zhao, Sida Peng, Hengkai Guo, Xiaowei Zhou, Guang Shi, Jiashi Feng, Bingyi Kang
Affiliation
Zhejiang University
Venue
ICLR 2026
Track
Oral

TL;DR

Depth Anything 3 uses a single vanilla DINOv2 transformer to take arbitrary input views and outputs consistent depth and ray maps, delivering leading pose, geometry, and visual rendering performance.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

transformer

← All ICLR 2026 Oral papers · Browse the whole archive