VibeVoice: Expressive Podcast Generation with Next-Token Diffusion
ICLROral2026
TL;DR
VibeVoice can synthesize long-form speech for up to 90 minutes (in a 64K context window length) with a maximum of 4 speakers, capturing the authentic conversational vibe and surpassing open-source and proprietary dialogue models.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
diffusion speech