FFN Fusion: Rethinking Sequential Computation in Large Language Models

NeurIPSSpotlight2025

Authors
Akhiad Bercovich, Mohammed Dabbah, Omri Puny, Ido Galil, Amnon Geifman, Yonatan Geifman, Izhak Golan, Ehud Dov Karpas, Itay Levy, Zach Moshe, Najeeb Nabwani, Tomer Ronen, Itamar Schen, Ido Shahaf, Oren Tropp, Ran Zilberstein, Ran El-Yaniv
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

We introduce \textit{FFN Fusion}, an architectural optimization technique that reduces sequential computation in large language models by identifying and exploiting natural opportunities for paralleli…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model optimization

← All NeurIPS 2025 Spotlight papers · Browse the whole archive