LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning

ICLROral2026

Authors
Yuhao Wu, Yushi Bai, Zhiqiang Hu, Roy Ka-Wei Lee, Juanzi Li
Affiliation
Singapore University of Technology and Design
Venue
ICLR 2026
Track
Oral

TL;DR

Ultra-long generation by large language models (LLMs) is a widely demanded scenario, yet it remains a significant challenge due to their maximum generation length limit and overall quality degradation as sequence length increases. Previous approaches, exemplified by LongWriter, typically rely on ''teaching'', which involves supervised fine-tunin...

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning large language model language model llm

← All ICLR 2026 Oral papers · Browse the whole archive