ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
NeurIPSOral2025
TL;DR
Multimodal large language models (MLLMs) extend LLMs to handle images, videos, and audio by incorporating feature extractors and projection modules. However, these additional components—combined with complex inference pipelines and heterogeneous workloads—introduce significant inference overhead.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model multimodal efficient audio video llm