Knowledge & Data · Software component

Unified Embedding Retriever

Software componentKnowledge & DataKnowledge & Dataarc:UnifiedEmbeddingRetriever

A multimodal retriever that encodes the text query into a shared text-image embedding space and retrieves text passages and images by cosine similarity from one store.

Responsibility. Retrieves text and images directly by cross-modal similarity in a unified space.

Also known as: Approach 1: Unified embedding space

Variant of Multimodal Retriever abstract

When to choose. Choose for general imagery, rapid prototyping, or minimal change to an existing text RAG stack with a single vector store.

sends data toinvokesis target of alternativeTois target of alternativeTospecializesContext Assembler: sends data toContext AssemblerJoint Multimodal Embedding Service: invokesJoint Multimodal Embeddi…Per-Modality Fan-Out Retriever: is target of alternativeToPer-Modality Fan-Out Ret…Grounded Text Retriever: is target of alternativeToGrounded Text RetrieverMultimodal Retriever: specializesMultimodal Retriever
Direct neighbourhood (hover for relationship types)

Relationships

invokes dependency

sends data to dynamic

alternative to variability

Classification

Patterns
Unified embedding space multimodal RAG
Technologies
CLIPNV Embed
Quality attributes
Maintainability (ISO/IEC 25010)Cost efficiency

Sources

  1. Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.