Knowledge & Data · Software component

Multimodal Context Assembler

Software componentKnowledge & DataKnowledge & Dataarc:MultimodalContextAssembler

A context assembler that detects retrieved chunks derived from images via metadata flags, loads the original images, and builds a multimodal prompt combining query, text context, and images.

Responsibility. Builds multimodal prompt context by re-attaching original images to image-derived chunks.

Also known as: Multimodal prompt builder

Variant of Context Assembler abstract

specializesreceives data fromis configured byreceives data fromreadssends data toContext Assembler: specializesContext AssemblerGrounded Text Retriever: receives data fromGrounded Text RetrieverMultimodal Chunk Metadata Schema: is configured byMultimodal Chunk Metadat…Cross-Modal Reranker: receives data fromCross-Modal RerankerSource Media Store: readsSource Media StoreMultimodal Answer Synthesizer: sends data toMultimodal Answer Synthe…
Direct neighbourhood (hover for relationship types)

Relationships

is configured by structural

reads dependency

receives data from dynamic

sends data to dynamic

Design guidance

Classification

Patterns
Retrieve on text, reason on original
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
Risks mitigated
Nuance lost when charts are compressed into text descriptions

Sources

  1. Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.