ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

NeurIPSSpotlight2025

Authors
Junliang Ye, Zhengyi Wang, Ruowen Zhao, Shenghao Xie, Jun Zhu
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Recently, the powerful text-to-image capabilities of GPT-4o have led to growing appreciation for native multimodal large language models…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model text-to-image multimodal llm 3d

← All NeurIPS 2025 Spotlight papers · Browse the whole archive