Model Serving · Software component
Fallback LLM Inference Service
Software componentModel ServingModelsarc:FallbackLLMInferenceService
A secondary LLM inference service from an independent provider, held in reserve to serve requests when the primary inference service is persistently degraded or quota-exhausted.
Responsibility. Serves LLM requests in place of a degraded primary inference service.
Also known as: Secondary LLM provider, Backup model
Relationships
is routed to by dynamic
is failover for control
Design guidance
- SHOULD run on infrastructure independent of the primary provider; shared failure domains render the fallback useless.
- SHOULD flag outputs served by the fallback with a degradation warning.
- SHOULD back high-availability deployments spread across multiple zones/regions with automatic failover to backup models.
Classification
- Patterns
- Provider failoverTransparent failover
- Technologies
- Anthropic Claude 3.5 Sonnet
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Single-provider outagePrimary quota exhaustion
Sources
- Ch2.8: T. Nguyen, "Error Handling and Resilience," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.8. ISBN: 9798244538229.
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.06: "Error Troubleshooting and Incident Response for Agent Systems," unpublished reference note (06-Error-Troubleshooting-Incident-Response.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note