MokA: Multimodal Low-Rank Adaptation for MLLMs
NeurIPSOral2025
TL;DR
In this paper, we reveal that most current efficient multimodal fine-tuning methods are hindered by a key limitation: they are directly borrowed from LLMs, often neglecting the intrinsic differences of multimodal scenarios and even affecting the full utilization of all modalities. Inspired by our empirical observation, we argue that unimodal adapta…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
fine-tuning multimodal efficient llm