Mixture-of-Experts Can Surpass Dense LLMs Under Strictly Equal Resource

ICLROral2026

Authors
Houyi Li, Ka Man Lo, Shijie Xuyang, Ziqi Wang, Wenzhen Zheng, Haocheng Zhang, Zhao Li, Shuigeng Zhou, Xiangyu Zhang, Daxin Jiang
Affiliation
Fudan University
Venue
ICLR 2026
Track
Oral

TL;DR

Mixture-of-Experts (MoE) language models dramatically expand model capacity and achieve remarkable performance without increasing per-token compute. However, can MoEs surpass dense architectures under strictly equal resource constraints — that is, when the total parameter count, training compute, and data budget are identical?

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model llm

← All ICLR 2026 Oral papers · Browse the whole archive