PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 27, 20240 citationsOpen Access

Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP

View Full Paper
RAReza AbbasiMSMohammad Mahdi SamieiMRMohammad Hossein Rohban

Key Points

Key points are not available for this paper at this time.

Abstract

Vision-language models, such as CLIP, have shown promising Out-of-Distribution (OoD) generalization under various types of distribution shifts. Recent studies attempted to investigate the leading cause of this capability. In this work, we follow the same path, but focus on a specific type of OoD data - images with novel compositions of attribute-object pairs - and study whether such models can successfully classify those images into composition classes. We carefully designed an authentic image test dataset called ImageNet-AO, consisting of attributes for objects that are unlikely encountered in the CLIP training sets. We found that CLIPs trained with large datasets such as OpenAI CLIP, LAION-400M, and LAION-2B show orders-of-magnitude improvement in effective compositional OoD generalization compared to both supervised models and CLIPs trained with smaller datasets, such as CC-12M and YFCC-15M. Our results provide evidence that the scale and diversity of training data and language supervision play a key role in unlocking the compositional generalization abilities of vision-language models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Abbasi et al. (2024) studied this question.

synapsesocial.com/papers/68e7230db6db64358769d3cehttps://doi.org/10.48550/arxiv.2403.18525
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Deciphering the Role of Representation Disentanglement: Investigating Compositional Generalization in CLIP Models2024
  2. 2Do CLIPs Always Generalize Better than ImageNet Models?2024
  3. 3CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models2024
  4. 4Evaluating Compositional Generalisation in VLMs and Diffusion Models2025
  5. 5Generalization Beyond Data Imbalance: A Controlled Study on CLIP for Transferable Insights2024