Tensor Product Attention Is All You Need
NeurIPSSpotlight2025
TL;DR
Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during inference…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model attention memory
← All NeurIPS 2025 Spotlight papers · Browse the whole archive