Perception Encoder: The best visual embeddings are not at the output of the network

NeurIPSOral2025

Authors
Daniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei, Tengyu Ma, Jiale Zhi, Jathushan Rajasegaran, Hanoona Abdul Rasheed, Junke Wang, Marco Monteiro, Hu Xu, Shiyu Dong, Nikhila Ravi, Shang-Wen Li, Piotr Dollar, Christoph Feichtenhofer
Affiliation
Meta
Venue
NeurIPS 2025
Track
Oral

TL;DR

We develop a CLIP model that is SotA on both image and video zero-shot recognition. Using its strong, general features we further create SotA encoders for language and spatial tasks.

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

zero-shot video

← All NeurIPS 2025 Oral papers · Browse the whole archive