Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think

NeurIPSOral2025

Authors
Ge Wu, Shen Zhang, Ruijing Shi, Shanghua Gao, Zhenyuan Chen, Lei Wang, Zhaowei Chen, Hongcheng Gao, Yao Tang, jian Yang, Ming-Ming Cheng, Xiang Li
Affiliation
Nankai University
Venue
NeurIPS 2025
Track
Oral

TL;DR

REPA and its variants effectively mitigate training challenges in diffusion models by incorporating external visual representations from pretrained models, through alignment between the noisy hidden projections of denoising networks and foundational clean image representations. We argue that the external alignment, which is absent during the entire…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

transformer alignment diffusion

← All NeurIPS 2025 Oral papers · Browse the whole archive