JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
NeurIPSSpotlight2025
TL;DR
This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model multimodal audio video llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive