HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models

NeurIPSOral2025

Authors
Zelin Peng, Zhengqin Xu, Qingyang Liu, Xiaokang Yang, Wei Shen
Affiliation
Shanghai Jiaotong University
Venue
NeurIPS 2025
Track
Oral

TL;DR

Multi-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computational resources (e.g., thousands of GPUs) for training to achieve cross-modal alignment at multi-granularity levels.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model alignment efficient llm

← All NeurIPS 2025 Oral papers · Browse the whole archive