Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction–Reasoning Synergy
ICLROral2026
TL;DR
Recent developments in Multimodal Large Language Models (MLLMs) have significantly improved Vision–Language (VL) reasoning in 2D domains. However, extending these capabilities to 3D scene understanding remains a major challenge.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model multimodal reasoning video llm 3d