LaViDa: A Large Diffusion Language Model for Multimodal Understanding

NeurIPSSpotlight2025

Authors
Shufan Li, Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Jason Kuen, Zhe Lin, Kai-Wei Chang, Aditya Grover
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

vision-language language model multimodal diffusion reasoning

← All NeurIPS 2025 Spotlight papers · Browse the whole archive