Model Serving · Software component
Model Router
Software componentModel ServingModelsarc:ModelRouter
A routing component that directs each query to a model tier according to predicted complexity, user tier or heuristics to minimise cost at acceptable quality.
Responsibility. Routes each query to the cheapest model tier able to answer it.
Also known as: Intelligent router, Complexity-based router, LLM fallback router, Rule-based model router, Provider failover router, Router Model, Complexity-Based Model Router, Router-first design, Smart router, Small-model routing, Specialty-based model selection, AI gateway, AI gateway routing layer, Workload-type configuration router, Model routing, Right-sizing model capacity
Relationships
is configured by structural
invokes dependency
is invoked by dependency
- Agent Controller abstract Ch1.8 Ch3.10 +2
reads dependency
- Response Cache abstract Ch2.8
- Secrets Vault Ch4.2
emits telemetry to dynamic
routes to dynamic
Design guidance
- SHOULD escalate complex queries to more capable models to preserve quality.
- MAY use a small embedding-based classifier to improve routing accuracy over heuristics.
- SHOULD attempt providers in a defined priority order until one succeeds, falling back only after retries on the primary are exhausted.
- SHOULD log the fallback level used for every request so operators can see how often fallbacks trigger and scale primary capacity.
- SHOULD account for changed cost and output quality when failing over transparently to a secondary model.
- SHOULD route simple queries (fact lookup, account inquiries, templates) to small models and multi-step reasoning or ambiguity resolution to large models.
- MAY default to small models and let users explicitly request detailed analysis to trigger large-model routing.
- SHOULD route straightforward queries to small efficient models and escalate to frontier models only when complexity requiring advanced reasoning is detected.
- SHOULD right-size model selection per task category rather than using one model universally.
- SHOULD implement model routing, cost caps and provider fallback declaratively once at the AI gateway rather than embedding model selection in every agent service.
- SHOULD centralise provider credentials and observability of all model interactions at the AI gateway.
- SHOULD route each request to the model configuration matching its workload class (interactive vs batch analytics).
- SHOULD route simple tasks to small models on low-tier hardware and complex reasoning to large models on top-tier GPUs, subject to latency SLAs.
- SHOULD route simple queries to cheap models using CoT and reserve ToT with expensive reasoning models for complex queries.
- SHOULD route requests across model variants by complexity or latency budget, with fallback, to trade quality against cost and enable A/B testing.
- SHOULD default moderate or ambiguous queries to the larger model, erring toward over-provisioning to avoid quality degradation.
- SHOULD validate per-tier quality on the routed query class before rollout.
Quantitative guidance
As stated by the sources; verify before use.
- 80/15/5 routing across small/standard/large tiers reduces cost 40-60% vs. all-large (Ch1.8).
- Reactive fallback adds typically 2-5 s of timeout-detection latency per request (Ch2.8).
- Fallbacks suit persistent outages longer than ~1 minute (Ch2.8).
- Example cost difference between primary and secondary models: $15 vs $3 per million tokens (Ch2.8).
- 70-80% of queries handled by small models, 20-30% by large models; blended cost ~$0.02/query vs $0.002 small and $0.08 large (Ch3.4).
- About 80% of production queries are straightforward and 20% need sophisticated reasoning (power-law distribution) (Ch3.10).
- Loan processing: small-model pre-screener handled 78% of volume; routing reduced cost more than universal use of either model (Ch3.10).
- E-commerce support routed FAQ questions to a 10x cheaper model (Ch3.10).
- Larger models cost 3-10x more while handling only marginally more complex tasks (Ch3.10).
- Example policy: queries under 500 tokens to a small fast model (Claude 3.5 Haiku); multi-step reasoning over 2,000 tokens to a stronger model (Claude 3.5 Sonnet) (Ch4.2).
- Query mix 30% simple / 50% moderate / 20% complex; routing the simple 30% to the small model cut monthly cost $3,930 -> $2,700 ($1,230) (Ch8.3).
- Small model matched the large model on simple queries: 99% task completion, 98% accuracy, 4.5/5 CSAT (Ch8.3).
- Typical savings of 40-60% from appropriate model sizing (Ref8.05).
Classification
- Patterns
- Ensemble scalingCascade by complexityHeuristic routing (length, keywords, user tier)Fallback chainTransparent failoverRule-based model routingHybrid small/large model routingConservative default-to-small with user-requested escalationRouter-first designDynamic model selection by task complexityEscalate-on-complexityCentralised declarative model routingMulti-provider fallbackHardware-tier routing (right-sizing)Small fast model first with fallback to larger model (cold-start mitigation)Model tiering: CoT on cheap models, ToT on reasoning models
- Technologies
- OpenAI GPT-4oAnthropic Claude 3.5 Sonnet
- Quality attributes
- Cost efficiencyFunctional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Transparency and accountability (NIST AI RMF: accountable and transparent)
- Risks mitigated
- Overspending on large models for simple queriesTotal failure during persistent provider outage or quota exhaustionMisrouting complex queries to small modelsBudget waste from excessive large-model routingRouting every query through expensive frontier modelsModel-selection logic duplicated and drifting across agent servicesRunaway spend when query volume spikesProvider outages
Sources
- Ch1.8: T. Nguyen, "Scalability and Production Deployment," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 1.8. ISBN: 9798244538229.
- Ch2.8: T. Nguyen, "Error Handling and Resilience," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 2.8. ISBN: 9798244538229.
- Ch3.4: T. Nguyen, "Tuning Model Parameters for Production Performance," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.4. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch4.5: T. Nguyen, "NVIDIA NIM and Triton Inference Server," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.5. ISBN: 9798244538229.
- Ch4.7: T. Nguyen, "Scaling Strategies," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.7. ISBN: 9798244538229.
- Ch5.2: T. Nguyen, "Tree-of-Thought (ToT) Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.2. ISBN: 9798244538229.
- Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
- Ref7.17: "Scaling Agentic AI Systems: Patterns and Strategies," unpublished reference note (17-Scalability-Patterns.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note