Advancing Expert Specialization for Better MoE
NeurIPSOral2025
TL;DR
Our proposed orthogonality and variance losses improve performance in downstream fine-tuning of Mixture-of-Experts models by enhancing expert specificity, addressing expert homogenization caused by load balancing, while maintaining load balance.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
fine-tuning