Optimal Sparsity of Mixture-of-Experts Language Models for Reasoning Tasks

ICLROral2026

Authors
Taishi Nakamura, Satoki Ishikawa, Masaki Kawamura, Takumi Okamoto, Daisuke Nohara, Jun Suzuki, Rio Yokota
Affiliation
Institute of Science Tokyo
Venue
ICLR 2026
Track
Oral

TL;DR

Memorization skills consistently benefit from higher sparsity, while reasoning skills require balancing active FLOPs with total tokens per parameter; the optimal point shifts with the compute budget.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model reasoning sparsity

← All ICLR 2026 Oral papers · Browse the whole archive