A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
NeurIPSSpotlight2025
TL;DR
Training high-performing Small Language Models (SLMs) remains computationally expensive, even with knowledge distillation and pruning from larger teacher models…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
language model distillation efficient
← All NeurIPS 2025 Spotlight papers · Browse the whole archive