Knowledge & Data · Software component
Multimodal Content Router
Software componentKnowledge & DataKnowledge & Dataarc:MultimodalContentRouter
A preprocessing router that dispatches each extracted content element to the processor suited to its modality and image type: chart extractor, image captioner, speech transcriber, or text chunker.
Responsibility. Selects the specialised processing path for each content element based on its detected modality and image class.
Also known as: Image routing logic, Modality detection and routing, Multimodal input handler, Input modality dispatcher
Relationships
invokes dependency
- Image Type Classifier abstract Ch2.7
is invoked by dependency
receives data from dynamic
routes to dynamic
is orchestrated by control
Design guidance
- SHOULD route information-dense images (charts, plots, graphs, tables, numeric diagrams) to structured chart extraction to preserve precise values.
- SHOULD route general images (photographs, illustrations, simple diagrams) to detailed captioning.
- MAY route complex diagrams with quantitative elements through both captioning and chart extraction.
- MAY use a joint image-text embedding as a baseline for rapid prototyping across diverse image types.
Classification
- Patterns
- Content-based routingDecision-tree routingHeterogeneous processing
- Technologies
- LlamaIndex MultiModal
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Cost efficiency
- Risks mitigated
- Loss of numerical precision when charts are processed by general-purpose modelsOver-complicated extraction of simple photos
Sources
- Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.
- Ch7.5: T. Nguyen, "NeMo Curator, Riva Speech AI & Multimodal Integration," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.5. ISBN: 9798244538229.