Enhancing CLIP Robustness via Cross-Modality Alignment
NeurIPSSpotlight2025
TL;DR
Vision-language models (VLMs) such as CLIP demonstrate strong generalization in zero-shot classification but remain highly vulnerable to adversarial perturbations…
Opening excerpt from the authors’ abstract. source
Read the paper
Topics
vision-language generalization language model adversarial robustness alignment zero-shot
← All NeurIPS 2025 Spotlight papers · Browse the whole archive