Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
NeurIPSSpotlight2025
TL;DR
Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted camera…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
vision-language language model
← All NeurIPS 2025 Spotlight papers · Browse the whole archive