NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

ICLROral2026

Authors
Chunrui Han, Guopeng Li, Jingwei Wu, Quan Sun, Yan Cai, Yuang Peng, Zheng Ge, Deyu Zhou, Haomiao Tang, Hongyu Zhou, Kenkun Liu, Shu-Tao Xia, Binxing Jiao, Daxin Jiang, Xiangyu Zhang, Yibo Zhu
Affiliation
Stepfun
Venue
ICLR 2026
Track
Oral

TL;DR

Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ vector quantization (VQ) to obtain discrete tokens with quantization loss. In this paper, we push the autoregressive paradigm forward with NextStep-1, a 14B autoregressi...

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

image generation text-to-image quantization diffusion

← All ICLR 2026 Oral papers · Browse the whole archive