Meta CLIP 2: A Worldwide Scaling Recipe
NeurIPSSpotlight2025
TL;DR
Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs)
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
large language model language model contrastive pretraining multimodal retrieval zero-shot llm
← All NeurIPS 2025 Spotlight papers · Browse the whole archive