Knowledge & Data · Software component
Unified Embedding Retriever
Software componentKnowledge & DataKnowledge & Dataarc:UnifiedEmbeddingRetriever
A multimodal retriever that encodes the text query into a shared text-image embedding space and retrieves text passages and images by cosine similarity from one store.
Responsibility. Retrieves text and images directly by cross-modal similarity in a unified space.
Also known as: Approach 1: Unified embedding space
Variant of Multimodal Retriever abstract
When to choose. Choose for general imagery, rapid prototyping, or minimal change to an existing text RAG stack with a single vector store.
Relationships
Classification
- Patterns
- Unified embedding space multimodal RAG
- Technologies
- CLIPNV Embed
- Quality attributes
- Maintainability (ISO/IEC 25010)Cost efficiency
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.