Caption This, Reason That: VLMs Caught in the Middle

NeurIPSSpotlight2025

Authors
Zihan Weng, Lucas Gomez, Taylor Whittington Webb, Pouya Bashivan
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Vision-Language Models (VLMs) have shown remarkable progress in visual understanding in recent years…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

vision-language language model

← All NeurIPS 2025 Spotlight papers · Browse the whole archive