Model Serving · Software component
Query Complexity Assessor
Software componentModel ServingModelsVariation point (abstract)arc:QueryComplexityAssessor
An abstract component that classifies an incoming query's complexity (e.g., simple, moderate, complex) before generation so that a router can select an appropriately sized model.
Responsibility. Classifies query complexity ahead of model selection.
Also known as: Complexity classification, estimate_query_complexity
Variants
| Variant | When to choose |
|---|---|
| Query Complexity Classifier | Choose when query phrasing varies widely and higher classification accuracy justifies the added latency and cost of a classifier model. |
| Rule-Based Complexity Classifier | Choose when query classes follow clear lexical patterns and zero added latency and cost matter more than robustness to phrasing variation. |
Relationships
is invoked by dependency
Design guidance
- SHOULD run before generation and add as little latency and cost as possible.
Quantitative guidance
As stated by the sources; verify before use.
- Complexity classes: simple (balances, history, definitions), moderate (performance summaries, holdings analysis), complex (tax optimization, multi-objective rebalancing, scenario analysis) (Ch8.3).
Classification
- Patterns
- Complexity-based model routing
- Quality attributes
- Performance efficiency (ISO/IEC 25010)
- Risks mitigated
- Over-provisioning large models for simple queriesQuality loss from under-provisioning complex queries
Sources
- Ch8.3: T. Nguyen, "Token Economics and Architecture," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.3. ISBN: 9798244538229.
- Ref8.05: "Cost Optimization and Resource Monitoring for Agent Systems," unpublished reference note (05-Cost-Optimization-Resource-Monitoring.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note