Unleashing Hour-Scale Video Training for Long Video-Language Understanding
NeurIPSSpotlight2025
TL;DR
Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs)
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
multimodal benchmark video
← All NeurIPS 2025 Spotlight papers · Browse the whole archive