Extrapolation by Association: Length Generalization Transfer In Transformers

NeurIPSSpotlight2025

Authors
Ziyang Cai, Nayoung Lee, Avi Schwarzschild, Samet Oymak, Dimitris Papailiopoulos
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Transformer language models have demonstrated impressive generalization capabilities in natural language domains, yet we lack a fine-grained understanding of how such generalization arises…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

generalization language model transformer

← All NeurIPS 2025 Spotlight papers · Browse the whole archive