Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
NeurIPSOral2025
TL;DR
REPA and its variants effectively mitigate training challenges in diffusion models by incorporating external visual representations from pretrained models, through alignment between the noisy hidden projections of denoising networks and foundational clean image representations. We argue that the external alignment, which is absent during the entire…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
transformer alignment diffusion