MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chains
ICLROral2026
TL;DR
With the increasing demand for step-wise, cross-modal, and knowledge-grounded reasoning, multimodal large language models (MLLMs) are evolving beyond the traditional fixed retrieve-then-generate paradigm toward more sophisticated agentic multimodal retrieval-augmented generation (MM-RAG). Existing benchmarks, however, mainly focus on simplified...
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model multimodal benchmark reasoning retrieval agent llm rag