EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

ICLROral2026

Authors
Dingdong WANG, Shujie LIU, Tianhua Zhang, Youjun Chen, Jinyu Li, Helen M. Meng
Affiliation
Chinese University of Hong Kong, The Chinese University of Hong Kong
Venue
ICLR 2026
Track
Oral

TL;DR

Emotional information in speech plays a unique role in multimodal perception. However, current Speech Large Language Models (SpeechLLMs), similar to conventional speech emotion recognition (SER) systems, still treat emotion understanding as a simple classification problem.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning large language model language model multimodal reasoning speech llm

← All ICLR 2026 Oral papers · Browse the whole archive