UMoE: Unifying Attention and FFN with Shared Experts
NeurIPSSpotlight2025
TL;DR
Sparse Mixture of Experts (MoE) architectures have emerged as a promising approach for scaling Transformer models…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
mixture of experts transformer attention
← All NeurIPS 2025 Spotlight papers · Browse the whole archive