Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource
ICLROral2026
TL;DR
Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token compute. However, can MoEs surpass dense architectures under strictly equal resource constraints — that is, when the total parameter count, training compute, and data budget are identical?
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model llm