FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving

NeurIPSSpotlight2025

Authors
Shuang Zeng, Xinyuan Chang, Mengwei Xie, Xinran Liu, Yifan Bai, Zheng Pan, Mu Xu, Xing Wei
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Vision–Language–Action (VLA) models are increasingly used for end-to-end driving due to their world knowledge and reasoning ability…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

reasoning

← All NeurIPS 2025 Spotlight papers · Browse the whole archive