NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
ICLROral2026
TL;DR
Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ vector quantization (VQ) to obtain discrete tokens with quantization loss. In this paper, we push the autoregressive paradigm forward with NextStep-1, a 14B autoregressi...
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
image generation text-to-image quantization diffusion