SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
NeurIPSSpotlight2025
TL;DR
KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
efficient llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive