Unleashing Hour-Scale Video Training for Long Video-Language Understanding

NeurIPSSpotlight2025

Authors
Jingyang Lin, Jialian Wu, Ximeng Sun, Ze Wang, Jiang Liu, Yusheng Su, Xiaodong Yu, Hao Chen, Jiebo Luo, Zicheng Liu, Emad Barsoum
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Recent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs)

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

multimodal benchmark video

← All NeurIPS 2025 Spotlight papers · Browse the whole archive