UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

ICLROral2026

Authors
Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh, Arushi Goel, Chao-Han Huck Yang, Wenliang Dai, Zihan Liu, Hanrong Ye, Shinji Watanabe, Mohammad Shoeybi, Bryan Catanzaro, Rafael Valle, Wei Ping
Affiliation
CMU, Carnegie Mellon University
Venue
ICLR 2026
Track
Oral

TL;DR

This paper introduces UALM, an audio language model designed to unify audio understanding, generation, and reasoning…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model reasoning audio

← All ICLR 2026 Oral papers · Browse the whole archive