Knowledge & Data · Software component
Joint Multimodal Embedding Service
Software componentKnowledge & DataKnowledge & Dataarc:JointMultimodalEmbeddingService
An embedding service that encodes both text and images into one shared vector space with jointly trained encoders, enabling cross-modal similarity search without conversion.
Responsibility. Embeds text and images into a unified embedding space.
Also known as: Unified embedding space, Multimodal embedder
Variant of Embedding Service abstract
When to choose. Choose for rapid deployment on existing text RAG with mostly general imagery (photos, simple diagrams); avoid for information-dense charts needing precise values or OCR.
Relationships
invokes dependency
is invoked by dependency
writes dependency
- Vector Index Store abstract Ch2.7
is routed to by dynamic
alternative to variability
Design guidance
- SHOULD L2-normalise embeddings so cosine similarity scores are comparable across queries and images.
- SHOULD NOT be used alone for charts, detailed diagrams, or images with text requiring OCR.
Quantitative guidance
As stated by the sources; verify before use.
- CLIP encoders produce 512-dimensional vectors (Ch2.7).
- NV Embed outputs 512 or 768 dimensions and is 3x faster than standard transformers at batch sizes 32-128 (Ch2.7).
- Embedding 10,000 image-text pairs: 8 min with vanilla CLIP vs under 3 min with NV Embed on an A100 (Ch2.7).
Classification
- Patterns
- Unified embedding space multimodal RAGContrastive learningZero-shot classification
- Technologies
- CLIPOpenCLIPNV Embed
- Quality attributes
- Maintainability (ISO/IEC 25010)
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.