DepthLM: Metric Depth from Vision Language Models

ICLROral2026

Authors
Zhipeng Cai, Ching-Feng Yeh, Hu Xu, Zhuang Liu, Gregory P. Meyer, Xinjie Lei, Changsheng Zhao, Shang-Wen Li, Vikas Chandra, Yangyang Shi
Affiliation
Meta
Venue
ICLR 2026
Track
Oral

TL;DR

The first proof that VLMs can have expert model level depth estimation accuracy without architecture or loss change…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model

← All ICLR 2026 Oral papers · Browse the whole archive