Advancing Expert Specialization for Better MoE

NeurIPSOral2025

Authors
Hongcan Guo, Haolang Lu, Guoshun Nan, Bolun Chu, Jialin Zhuang, Yuan Yang, Wenhao Che, Xinye Cao, Sicong Leng, Qimei Cui, Xudong Jiang
Affiliation
ByteDance Inc.
Venue
NeurIPS 2025
Track
Oral

TL;DR

Our proposed orthogonality and variance losses improve performance in downstream fine-tuning of Mixture-of-Experts models by enhancing expert specificity, addressing expert homogenization caused by load balancing, while maintaining load balance.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

fine-tuning

← All NeurIPS 2025 Oral papers · Browse the whole archive