Observability & Evaluation · Software component
Online Evaluator
Software componentObservability & EvaluationObservability & Evaluationarc:OnlineEvaluator
An evaluation component that scores a sample of real production interactions without ground truth, combining judge-model scores, implicit behavioural signals, and explicit user feedback.
Responsibility. Continuously evaluates agent quality on live production traffic.
Also known as: Production evaluation, Online monitoring, Continuous production monitoring, Online evaluation, Continuous monitoring, Longitudinal monitoring, Continuous alignment monitoring, Drift monitor, Production hallucination tracking, Online Evaluation Monitor, Continuous canary quality evaluation, Per-query-type performance analytics, Continuous quality monitoring, Real-time agent scoring
Relationships
is configured by structural
invokes dependency
reads dependency
writes dependency
receives data from dynamic
sends data to dynamic
triggers dynamic
evaluates assurance
monitors assurance
Design guidance
- SHOULD combine implicit and explicit feedback, accounting for response and sampling bias in explicit ratings.
- SHOULD feed discovered failures back into offline test sets.
- MUST NOT treat offline benchmarks as the final validation gate; online monitoring with real user data is required.
- SHOULD alert when accuracy, latency, satisfaction or error-rate metrics shift beyond acceptable thresholds.
- SHOULD trigger retraining when performance degradation from drift is detected.
- SHOULD track correlation between reward-model predictions and current human judgments to catch misalignment.
- SHOULD log all outputs with query, retrieved documents, tool results, confidence scores and reasoning traces for retrospective analysis.
- SHOULD randomly sample production requests from both canary and stable versions and score them with an LLM judge.
- SHOULD record success, steps taken, hallucination, cost and latency per execution and compare them against baseline.
- SHOULD monitor at conversation, agent-trend and system-health levels.
Quantitative guidance
As stated by the sources; verify before use.
- Canary example samples 10% of conversations; canary correctness 0.82 vs stable 0.88 would trigger rollback (Ch4.2).
- Alert thresholds: success rate drops >5%, cost per request doubles, P95 exceeds SLA, hallucination rate >2% (Ref8.03).
Classification
- Patterns
- Online evaluationAdaptive samplingDrift-triggered retrainingContinuous learning
- Technologies
- NVIDIA NeMoGalileo
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Cost efficiencyReliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Model driftDistribution shiftSource document driftUser expectation mismatchData drift degrading fine-tuned modelsReward model distribution shiftReward hackingDistribution collapseDiscovering failures only through user complaintsUndetected production degradation
Sources
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ch6.6: T. Nguyen, "Query Decomposition and Adaptive Retrieval," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.6. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.
- Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
- Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref10.05: "The Data Flywheel: Continuous Improvement Loop," unpublished reference note (05-Data-Flywheel-Continuous-Improvement.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note