LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
ICLROral2026
TL;DR
Ultra-long generation by large language models (LLMs) is a widely demanded scenario, yet it remains a significant challenge due to their maximum generation length limit and overall quality degradation as sequence length increases. Previous approaches, exemplified by LongWriter, typically rely on ''teaching'', which involves supervised fine-tunin...
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reinforcement learning large language model language model llm