Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

NeurIPSSpotlight2025

Authors
Diankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi Duan
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model multimodal llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive