Knowledge & Data · Model asset
Contrastive Image-Text Encoder
Model assetKnowledge & DataKnowledge & Dataarc:ContrastiveImageTextEncoder
A pair of text and image encoders trained jointly with contrastive loss so matching text-image pairs map to nearby vectors in one shared space.
Responsibility. Maps text and images into a shared embedding space.
Relationships
deployed on structural
is optimized by lifecycle
Quantitative guidance
As stated by the sources; verify before use.
- CLIP produces 512-dimensional text and image vectors (Ch2.7).
Classification
- Patterns
- Contrastive learningZero-shot classification
- Technologies
- CLIPOpenCLIPNV Embed
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.