Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs

NeurIPSSpotlight2025

Authors
Insu Lee, Wooje Park, Jaeyun Jang, Minyoung Noh, Kyuhong Shim, Byonghyo Shim
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Large vision-language models (LVLMs) are increasingly deployed in interactive applications such as virtual and augmented reality, where a first-person (egocentric) view captured by head-mounted camera…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

vision-language language model

← All NeurIPS 2025 Spotlight papers · Browse the whole archive