Let Features Decide Their Own Solvers: Hybrid Feature Caching for Diffusion Transformers
ICLROral2026
TL;DR
Diffusion Transformers offer state-of-the-art fidelity in image and video synthesis, but their iterative sampling process remains a major bottleneck due to the high cost of transformer forward passes at each timestep. To mitigate this, feature caching has emerged as a training-free acceleration technique that reuses hidden representations.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
transformer diffusion video