ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models

ICLROral2026

Authors
Akshat Ramachandran, Marina Neseem, Charbel Sakr, Rangharajan Venkatesan, Brucek Khailany, Tushar Krishna
Affiliation
Georgia Institute of Technology
Venue
ICLR 2026
Track
Oral

TL;DR

The long-output context generation of large reasoning models enables extended chain of thought (CoT) but also drives rapid growth of the key–value (KV) cache, quickly overwhelming GPU memory. To address this challenge, we propose ThinKV, a thought-adaptive KV cache compression framework.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

efficient reasoning memory

← All ICLR 2026 Oral papers · Browse the whole archive