Observability & Evaluation · Software component

Online Evaluator

Software componentObservability & EvaluationObservability & Evaluationarc:OnlineEvaluator

An evaluation component that scores a sample of real production interactions without ground truth, combining judge-model scores, implicit behavioural signals, and explicit user feedback.

Responsibility. Continuously evaluates agent quality on live production traffic.

Also known as: Production evaluation, Online monitoring, Continuous production monitoring, Online evaluation, Continuous monitoring, Longitudinal monitoring, Continuous alignment monitoring, Drift monitor, Production hallucination tracking, Online Evaluation Monitor, Continuous canary quality evaluation, Per-query-type performance analytics, Continuous quality monitoring, Real-time agent scoring

monitors; evaluatesevaluatesmonitorstriggerssends data toreadsmonitorsinvokesreceives data fromtriggersreadsevaluateswritessends data tosends data toreadsmonitorsreceives data fromFine-Tuned Agent Model: monitors; evaluatesFine-Tuned Agent ModelAgent Controller: evaluatesAgent ControllerLLM Inference Service: monitorsLLM Inference ServiceAlert Manager: triggersAlert ManagerMetrics Collector: sends data toMetrics CollectorTrace Store: readsTrace StoreAnswer Synthesizer: monitorsAnswer SynthesizerLLM Judge: invokesLLM JudgeFeedback Collector: receives data fromFeedback CollectorFine-Tuning Pipeline: triggersFine-Tuning PipelineUser Feedback Store: readsUser Feedback StoreDialogue Flow Manager: evaluatesDialogue Flow ManagerEvaluation Dataset: writesEvaluation DatasetExperiment Guardrail Monitor: sends data toExperiment Guardrail Mon…Quality Drift Detector: sends data toQuality Drift DetectorEvaluation Baseline: readsEvaluation BaselineReward Model: monitorsReward ModelBehavioral Signal Tracker: receives data fromBehavioral Signal Tracker+7 more (see relationships)
Direct neighbourhood (hover for relationship types)

Relationships

is configured by structural

invokes dependency

reads dependency

writes dependency

receives data from dynamic

sends data to dynamic

triggers dynamic

evaluates assurance

monitors assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Online evaluationAdaptive samplingDrift-triggered retrainingContinuous learning
Technologies
NVIDIA NeMoGalileo
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Cost efficiencyReliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)Maintainability (ISO/IEC 25010)
Risks mitigated
Model driftDistribution shiftSource document driftUser expectation mismatchData drift degrading fine-tuned modelsReward model distribution shiftReward hackingDistribution collapseDiscovering failures only through user complaintsUndetected production degradation

Sources

  1. Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
  2. Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
  3. Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
  4. Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
  5. Ch6.6: T. Nguyen, "Query Decomposition and Adaptive Retrieval," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.6. ISBN: 9798244538229.
  6. Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
  7. Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.
  8. Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
  9. Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  10. Ref10.05: "The Data Flywheel: Continuous Improvement Loop," unpublished reference note (05-Data-Flywheel-Continuous-Improvement.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note