FlashVID: Efficient Video Large Language Models via Training-free Tree-based Spatiotemporal Token Merging
ICLROral2026
TL;DR
We introduce FlashVID, a training-free and plug-and-play inference acceleration framework for Video LLMs, enabling a satisfactory speedup with negligible performance degradation.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model efficient video llm