Tensor Product Attention Is All You Need

NeurIPSSpotlight2025

Authors
Yifan Zhang, Yifeng Liu, Huizhuo Yuan, Zhen Qin, Yang Yuan, Quanquan Gu, Andrew C Yao
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during inference…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

language model attention memory

← All NeurIPS 2025 Spotlight papers · Browse the whole archive