VibeVoice: Expressive Podcast Generation with Next-Token Diffusion

ICLROral2026

Authors
Zhiliang Peng, Jianwei Yu, Wenhui Wang, Yaoyao Chang, Yutao Sun, Li Dong, Yi Zhu, Weijiang Xu, Hangbo Bao, Zehua Wang, Shaohan Huang, Yan Xia, Furu Wei
Affiliation
Microsoft
Venue
ICLR 2026
Track
Oral

TL;DR

VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational vibe and surpassing open-source and proprietary dialogue models.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

diffusion speech

← All ICLR 2026 Oral papers · Browse the whole archive