AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders

NeurIPSSpotlight2025

Authors
Yuezhou Hu, Jiaxin Guo, Xinyu Feng, Tuo Zhao
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Speculative Decoding (SD) accelerates large language model inference by employing a small draft model to generate predictions, which are then verified by a larger target model…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model distillation efficient

← All NeurIPS 2025 Spotlight papers · Browse the whole archive