VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

NeurIPSSpotlight2025

Authors
Haozhe Wang, Chao Qu, Zuming Huang, Wei Chu, Fangzhen Lin, Wenhu Chen
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Recently, slow-thinking systems like GPT-o1 and DeepSeek-R1 have demonstrated great potential in solving challenging problems through explicit reflection…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reinforcement learning vision-language language model

← All NeurIPS 2025 Spotlight papers · Browse the whole archive