Visual Planning: Let's Think Only with Images

ICLROral2026

Authors
Yi Xu, Chengzu Li, Han Zhou, Xingchen Wan, Caiqi Zhang, Anna Korhonen, Ivan Vulić
Affiliation
Mistral AI
Venue
ICLR 2026
Track
Oral

TL;DR

Recent advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have substantially enhanced machine reasoning across diverse tasks. However, these models predominantly rely on pure text as the medium for both expressing and structuring reasoning, even when visual information is present.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

large language model language model multimodal reasoning planning llm

← All ICLR 2026 Oral papers · Browse the whole archive