FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging

ICLROral2026

Authors
Ziyang Fan, Keyu Chen, Ruilong Xing, Yulin Li, Li Jiang, Zhuotao Tian
Affiliation
Harbin Institute of Technology, Shenzhen
Venue
ICLR 2026
Track
Oral

TL;DR

We introduce FlashVID, a training-free and plug-and-play inference acceleration framework for Video LLMs, enabling a satisfactory speedup with negligible performance degradation.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model efficient video llm

← All ICLR 2026 Oral papers · Browse the whole archive