Knowledge & Data · Software component

Grounded Text Retriever

Software componentKnowledge & DataKnowledge & Dataarc:GroundedTextRetriever

A multimodal retriever that searches a text index containing native text plus text generated from images and audio, returning chunks with source-modality metadata.

Responsibility. Retrieves grounded-text chunks and their modality provenance by text similarity.

Also known as: Approach 2: Ground to text with metadata

Variant of Multimodal Retriever abstract

When to choose. Choose for information-dense visuals (financial reports, scientific charts, technical diagrams) where precise chart content must be retrievable and existing text RAG infrastructure should stay unchanged.

invokessends data toalternative tospecializesalternative toText Embedding Service: invokesText Embedding ServiceMultimodal Context Assembler: sends data toMultimodal Context Assem…Per-Modality Fan-Out Retriever: alternative toPer-Modality Fan-Out Ret…Multimodal Retriever: specializesMultimodal RetrieverUnified Embedding Retriever: alternative toUnified Embedding Retrie…
Direct neighbourhood (hover for relationship types)

Relationships

invokes dependency

sends data to dynamic

alternative to variability

Classification

Patterns
Ground-to-text multimodal RAGRetrieve on text, reason on original
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)
Risks mitigated
Undiscoverable numeric values inside charts

Sources

  1. Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.