Also known as: CLIP (Contrastive Language-Image Pretraining) Β· CLIP Β· Contrastive Language-Image Pretraining Β· Contrastive Language Image Pretraining (CLIP)
CLIP, or Contrastive Language-Image Pretraining, is a model approach that learns relationships between images and natural language descriptions. It enables tasks such as image-text matching, zero-shot image classification, and multimodal retrieval.