AdaSPEC: Selective Knowledge Distillation for Efficient Speculative Decoders
NeurIPSSpotlight2025
TL;DR
Speculative Decoding (SD) accelerates large language model inference by employing a small draft model to generate predictions, which are then verified by a larger target model…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model distillation efficient
← All NeurIPS 2025 Spotlight papers · Browse the whole archive