VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
NeurIPSSpotlight2025
TL;DR
Recent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model multimodal speech llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive