Latent Speech-Text Transformer

ICLROral2026

Authors
Yen-Ju Lu, Yashesh Gaur, Wei Zhou, Benjamin Muller, Jesus Villalba, Najim Dehak, Luke Zettlemoyer, Gargi Ghosh, Mike Lewis, Srini Iyer, Duc Le
Affiliation
Johns Hopkins University
Venue
ICLR 2026
Track
Oral

TL;DR

We introduce Latent Speech-Text Transformer, which patches long speech token sequences into latent units, improving text–speech transfer while cutting pre-training and inference compute, and significantly outperforming existing speech-text LLMs.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

pre-training transformer speech llm

← All ICLR 2026 Oral papers · Browse the whole archive