Knowledge & Data · Software component
Multimodal Context Assembler
Software componentKnowledge & DataKnowledge & Dataarc:MultimodalContextAssembler
A context assembler that detects retrieved chunks derived from images via metadata flags, loads the original images, and builds a multimodal prompt combining query, text context, and images.
Responsibility. Builds multimodal prompt context by re-attaching original images to image-derived chunks.
Also known as: Multimodal prompt builder
Variant of Context Assembler abstract
Relationships
is configured by structural
reads dependency
receives data from dynamic
sends data to dynamic
Design guidance
- SHOULD pass the original image with the query to a multimodal LLM when retrieved chunks originated from image captions.
Classification
- Patterns
- Retrieve on text, reason on original
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- Nuance lost when charts are compressed into text descriptions
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.