Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction–Reasoning Synergy

ICLROral2026

Authors
Haijier Chen, Bo Xu, Shoujian zhang, Haoze Liu, Jiaxuan Lin, Jingrong Wang
Affiliation
Wuhan University
Venue
ICLR 2026
Track
Oral

TL;DR

Recent developments in Multimodal Large Language Models (MLLMs) have significantly improved Vision–Language (VL) reasoning in 2D domains. However, extending these capabilities to 3D scene understanding remains a major challenge.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model multimodal reasoning video llm 3d

← All ICLR 2026 Oral papers · Browse the whole archive