Differentiable Hierarchical Visual Tokenization
NeurIPSSpotlight2025
TL;DR
Vision Transformers rely on fixed patch tokens that ignore the spatial and semantic structure of images…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
transformer
← All NeurIPS 2025 Spotlight papers · Browse the whole archive