UMoE: Unifying Attention and FFN with Shared Experts

NeurIPSSpotlight2025

Authors
Yuanhang Yang, Chaozheng Wang, Jing Li
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Sparse Mixture of Experts (MoE) architectures have emerged as a promising approach for scaling Transformer models…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

mixture of experts transformer attention

← All NeurIPS 2025 Spotlight papers · Browse the whole archive