Model Serving · Software component

Inference Engine Selector

Software componentModel ServingModelsarc:InferenceEngineSelector

A startup component that inspects GPU architecture, compute capability and VRAM, selects a pre-compiled engine matching the detected hardware, and falls back to a portable runtime when none exists.

Responsibility. Chooses the best available inference engine for the detected hardware.

Also known as: Hardware detection cascade, Automatic optimization

deployed onroutes toreadsroutes toInference Server: deployed onInference ServerPre-compiled Engine Backend: routes toPre-compiled Engine Back…Model Repository: readsModel RepositoryPortable LLM Runtime Backend: routes toPortable LLM Runtime Bac…
Direct neighbourhood (hover for relationship types)

Relationships

deployed on structural

reads dependency

routes to dynamic

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Hardware-aware engine selectionGraceful fallback
Technologies
NVIDIA NIM
Quality attributes
Performance efficiency (ISO/IEC 25010)Flexibility (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)

Sources

  1. Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.