Orchestration · Software component
RAG Query Orchestrator
Software componentOrchestrationOrchestration & Toolsarc:RAGQueryOrchestrator
A stateless query-time service that sequences cache lookup, retrieval, reranking, context assembly and grounded generation for each RAG request, recording per-stage timing, token usage and cache status.
Responsibility. Coordinates the per-request retrieval-augmented generation pipeline.
Also known as: RAG service, Production RAG service, Retrieval service, Query endpoint, Decomposed RAG pipeline, Hybrid RAG pipeline
Relationships
deployed on structural
exposes structural
invokes dependency
is cached by dependency
- Response Cache abstract Ch6.5
reads dependency
writes dependency
emits telemetry to dynamic
is routed to by dynamic
is constrained by control
is orchestrated by control
is scaled by control
- Autoscaler abstract Ch6.5
orchestrates control
is monitored by assurance
Design guidance
- MUST be stateless, storing no request-specific data locally, so load balancers can distribute requests to any replica.
- SHOULD check the response cache before any embedding, retrieval or generation work.
- SHOULD assign a unique request ID and record per-stage latency (retrieval vs. generation), token usage and cache-hit status in every response and log entry.
- SHOULD distinguish expected empty results (warn, continue) from dependency failures (return 503) so monitoring separates client from service errors.
- SHOULD initialize persistent, pooled connections at startup with short connect and longer read timeouts rather than per request.
Quantitative guidance
As stated by the sources; verify before use.
- Production targets: P50 < 1 s, P95 < 2 s, P99 < 5 s; stage targets embedding 50-100 ms, retrieval 100-300 ms, generation 500-1500 ms (Ch6.5).
- Throughput headroom: operate at 50-70% of maximum sustainable QPS; sustain 2x normal load (e.g., 20,000 vs 10,000 queries/hour) during spikes (Ch6.5).
- Per-query cost target $0.001-0.01 (embedding ~$0.0001, retrieval ~$0.0001, generation $0.001-0.01) (Ch6.5).
- Error rate target < 0.1%; uptime SLA 99.9% (8.76 h/yr) or 99.95% (4.38 h/yr) (Ch6.5).
- Example connection timeouts: 5 s connect, 60 s read (Ch6.5).
Classification
- Patterns
- Layered production RAG architectureCache-first architectureStateless service for horizontal scalingRequest-ID log correlationPersistent client connection pooling
- Technologies
- FastAPIPydanticAsyncOpenAI clientWeaviate client
- Quality attributes
- Performance efficiency (ISO/IEC 25010)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Cost efficiencyMaintainability (ISO/IEC 25010)
- Risks mitigated
- Connection exhaustion from per-request database connectionsAccepting traffic before dependencies are availableUndiagnosable latency without per-stage timing
Sources
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch6.6: T. Nguyen, "Query Decomposition and Adaptive Retrieval," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.6. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.