SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMs
NeurIPSSpotlight2025
TL;DR
Large language models (LLMs) power many modern applications, but serving them at scale remains costly and resource-intensive…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive