Differentiable Hierarchical Visual Tokenization

NeurIPSSpotlight2025

Authors
Marius Aasan, Martine Hjelkrem-Tan, Nico Catalano, Changkyu Choi, Adín Ramírez Rivera
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Vision Transformers rely on fixed patch tokens that ignore the spatial and semantic structure of images…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

transformer

← All NeurIPS 2025 Spotlight papers · Browse the whole archive