Perception Encoder: The best visual embeddings are not at the output of the network
NeurIPSOral2025
TL;DR
We develop a CLIP model that is SotA on both image and video zero-shot recognition. Using its strong, general features we further create SotA encoders for language and spatial tasks.
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
zero-shot video