Model Serving · Software component
Inference Engine Selector
Software componentModel ServingModelsarc:InferenceEngineSelector
A startup component that inspects GPU architecture, compute capability and VRAM, selects a pre-compiled engine matching the detected hardware, and falls back to a portable runtime when none exists.
Responsibility. Chooses the best available inference engine for the detected hardware.
Also known as: Hardware detection cascade, Automatic optimization
Relationships
deployed on structural
reads dependency
routes to dynamic
Quantitative guidance
As stated by the sources; verify before use.
- Manual engine compilation and tuning takes 10-15 hours per model; engine download is 3-5GB taking 2-3 minutes (Ch7.1B).
Classification
- Patterns
- Hardware-aware engine selectionGraceful fallback
- Technologies
- NVIDIA NIM
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Flexibility (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Sources
- Ch7.1B: T. Nguyen, "Nvidia NIM and Colang," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.1B. ISBN: 9798244538229.