Enhancing CLIP Robustness via Cross-Modality Alignment

NeurIPSSpotlight2025

Authors
Xingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao, Hanwang Zhang
Venue
NeurIPS 2025
Track
Spotlight

TL;DR

Vision-language models (VLMs) such as CLIP demonstrate strong generalization in zero-shot classification but remain highly vulnerable to adversarial perturbations…

Opening excerpt from the authors’ abstract. source

Read the paper

Topics

vision-language generalization language model adversarial robustness alignment zero-shot

← All NeurIPS 2025 Spotlight papers · Browse the whole archive