LaViDa: A Large Diffusion Language Model for Multimodal Understanding
NeurIPSSpotlight2025
TL;DR
Modern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
vision-language language model multimodal diffusion reasoning
← All NeurIPS 2025 Spotlight papers · Browse the whole archive