ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models

ICLROral2026

Authors
Federico Danieli, Pau Rodriguez, Miguel Sarabia, Xavier Suau, Luca Zappella
Affiliation
Apple
Venue
ICLR 2026
Track
Oral

TL;DR

We break the sequential bottleneck of nonlinear RNNs, enabling training of billion-scale LSTM/GRU models, competitive with modern architectures…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model

← All ICLR 2026 Oral papers · Browse the whole archive