FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
NeurIPSSpotlight2025
TL;DR
The rapid progress of large language models (LLMs) has catalyzed the emergence of multimodal large language models (MLLMs) that unify visual understanding and image generation within a single framewor…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model image generation language model multimodal llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive