FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities

NeurIPSSpotlight2025

Authors
Jin Wang, Yao Lai, Aoxue Li, Shifeng Zhang, Jiacheng Sun, Ning Kang, Chengyue Wu, Zhenguo Li, Ping Luo
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

The rapid progress of large language models (LLMs) has catalyzed the emergence of multimodal large language models (MLLMs) that unify visual understanding and image generation within a single framewor…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model image generation language model multimodal llm

← All NeurIPS 2025 Spotlight papers · Browse the whole archive