ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
ICLROral2026
TL;DR
The long-output context generation of large reasoning models enables extended chain of thought (CoT) but also drives rapid growth of the key–value (KV) cache, quickly overwhelming GPU memory. To address this challenge, we propose ThinKV, a thought-adaptive KV cache compression framework.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
efficient reasoning memory