SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMs

NeurIPSSpotlight2025

Authors
Jinwoo Park, Seunggeun Cho, Dongsu Han
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Large language models (LLMs) power many modern applications, but serving them at scale remains costly and resource-intensive…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive