Knowledge & Data · Software component

VLM-based Image Type Classifier

Software componentKnowledge & DataKnowledge & Dataarc:VLMImageTypeClassifier

An image type classifier that prompts a vision-language model to categorise each image (e.g., 'chart/plot' versus 'general') before routing.

Responsibility. Uses an existing vision-language model to classify images for routing.

Also known as: Meta-classification with the VLM

Variant of Image Type Classifier abstract

When to choose. Choose when routing at scale and a vision-language model is already deployed, reusing it for meta-classification instead of adding a separate classification model.

invokesspecializesis target of alternativeToOpenAI-Compatible Inference API: invokesOpenAI-Compatible Infere…Image Type Classifier: specializesImage Type ClassifierHeuristic Image Type Classifier: is target of alternativeToHeuristic Image Type Cla…
Direct neighbourhood (hover for relationship types)

Relationships

invokes dependency

alternative to variability

Design guidance

Classification

Patterns
Meta-classification
Technologies
NVIDIA NeVA 22B
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Maintainability (ISO/IEC 25010)

Sources

  1. Ch2.7: T. Nguyen, "Multimodal RAG Approaches," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.7. ISBN: 9798244538229.