ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism

NeurIPSOral2025

Authors
Zedong Liu, Shenggan Cheng, Guangming Tan, Yang You, Dingwen Tao
Affiliation
University of Electronic Science and Technology of China
Venue
NeurIPS 2025
Track
Oral

TL;DR

Multimodal large language models (MLLMs) extend LLMs to handle images, videos, and audio by incorporating feature extractors and projection modules. However, these additional components—combined with complex inference pipelines and heterogeneous workloads—introduce significant inference overhead.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model multimodal efficient audio video llm

← All NeurIPS 2025 Oral papers · Browse the whole archive