FFN Fusion: Rethinking Sequential Computation in Large Language Models
NeurIPSSpotlight2025
TL;DR
We introduce \textit{FFN Fusion}, an architectural optimization technique that reduces sequential computation in large language models by identifying and exploiting natural opportunities for paralleli…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model optimization
← All NeurIPS 2025 Spotlight papers · Browse the whole archive