HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
NeurIPSOral2025
TL;DR
Multi-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computational resources (e.g., thousands of GPUs) for training to achieve cross-modal alignment at multi-granularity levels.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model alignment efficient llm