ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
ICLROral2026
TL;DR
We break the sequential bottleneck of nonlinear RNNs, enabling training of billion-scale LSTM/GRU models, competitive with modern architectures…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model