StreamForest: Efficient Online Video Understanding with Persistent Event Memory

NeurIPSSpotlight2025

Authors
Xiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li, Jing Wang, Jiaxin Li, Ziang Yan, Kun Tian, Meng Tian, Xinhai Zhao, Yi Wang, Limin Wang
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model multimodal efficient memory video llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive