Model Serving · Software component
Edge Inference Runtime
Software componentModel ServingModelsarc:EdgeInferenceRuntime
An on-device inference runtime that executes compact, hardware-targeted model formats locally, operating autonomously when connectivity is absent.
Responsibility. Executes optimized models on edge hardware.
Also known as: On-device inference engine
Relationships
deployed on structural
- Edge Device abstract Ch4.3
hosts structural
is failover for control
is monitored by assurance
Quantitative guidance
As stated by the sources; verify before use.
- Worked example on Jetson Nano: YOLOv8-m 120 ms / 2.1 GB (FP32) -> 26 ms / 8.8 MB model with 3.4% mAP loss after INT8 PTQ, structured pruning and runtime optimization (Ch4.3).
Classification
- Patterns
- Offline inferenceLocal fallback
- Technologies
- TensorFlow LiteONNX RuntimePyTorch MobileIntel OpenVINOApple Core MLTensorRT
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Transparency and accountability (NIST AI RMF: accountable and transparent)
- Risks mitigated
- Loss of capability when offline
Sources
- Ch4.3: T. Nguyen, "Container Orchestration and Edge Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.3. ISBN: 9798244538229.