FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
NeurIPSSpotlight2025
TL;DR
Vision–Language–Action (VLA) models are increasingly used for end-to-end driving due to their world knowledge and reasoning ability…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
reasoning
← All NeurIPS 2025 Spotlight papers · Browse the whole archive