Visual Planning: Let's Think Only with Images
ICLROral2026
TL;DR
Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across diverse tasks. However, these models predominantly rely on pure text as the medium for both expressing and structuring reasoning, even when visual information is present.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model multimodal reasoning planning llm