SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference

NeurIPSSpotlight2025

Authors
Yi Zhao, Yajuan Peng, Nguyen Cam-Tu, Zuchao Li, Wang Xiaoliang, hai zhao, Xiaoming Fu
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

efficient llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive