Human Oversight · Human role
Human Evaluator
Human roleHuman OversightExperience & Human Oversightarc:HumanEvaluator
A domain expert who scores agent outputs on subjective criteria using structured rubrics, after calibration training and with full task context.
Responsibility. Provides rubric-based human judgment of agent output quality.
Also known as: Domain expert evaluator, Subject matter expert, Human-in-the-loop tester, Domain expert reviewer, Clinical evaluation team, Risk management reviewer, Human-in-the-loop evaluation, Blind evaluator, Shadow-mode reviewer, Quality assurance reviewer
Relationships
is configured by structural
invokes dependency
- Feedback Collector abstract Ch10.5
- Trace Annotation Console Ch3.9 Ref8.01 +1
is triggered by dynamic
receives data from dynamic
receives escalation from dynamic
sends data to dynamic
audits assurance
evaluates assurance
- Agent Controller abstract Ch3.3 Ch10.5
- Answer Synthesizer abstract Ch3.10 Ch6.5
- Fine-Tuned Agent Model Ch10.3
- Intent Router Ch10.1
- LLM Judge abstract Ch3.3 Ch3.9 +1
- Reasoning Engine Ch3.9
- Tree Search Controller abstract Ch5.2
Design guidance
- SHOULD receive calibration training and full context (prompt, retrieved information, intended use) before scoring.
- SHOULD evaluate blind to agent version, model or configuration, with randomized presentation order.
- SHOULD use multiple independent reviewers and calibration sessions before large-scale evaluation.
- SHOULD be targeted at gold-standard creation and sampled validation because expert time does not scale to continuous production monitoring.
- SHOULD evaluate blind to which agent version produced each output.
- SHOULD measure inter-rater agreement and refine rubrics when disagreement exceeds 20%.
- SHOULD review a 5-10% sample of production traces (random, error and edge-case samples) against a scorecard of completion, tool appropriateness, hallucination and helpfulness.
- SHOULD manually review high-reward-scoring responses to verify genuine quality rather than reward-artifact exploitation.
Quantitative guidance
As stated by the sources; verify before use.
- Rubric scales of 1-5 per criterion over 50-100 samples per category for statistical significance (Ch3.3).
- Inter-rater agreement SHOULD exceed 85% (Ref8.03).
Classification
- Patterns
- Rubric-based scoringEvaluator calibration trainingBlind evaluationMultiple independent reviewers with inter-rater reliabilityCalibration sessionsGold-standard dataset creationDefense-in-depth stage 5 (strategic sampling)
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- Unquantifiable quality criteria missed by automated metricsConfirmation bias in human assessmentHalo and recency effectsGround truth gapConfirmation bias in hallucination assessment
Sources
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch3.9: T. Nguyen, "Reasoning Quality," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.9. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch5.2: T. Nguyen, "Tree-of-Thought (ToT) Fundamentals," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.2. ISBN: 9798244538229.
- Ch6.5: T. Nguyen, "Production RAG Systems," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 6.5. ISBN: 9798244538229.
- Ch10.1: T. Nguyen, "Conversational UI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.1. ISBN: 9798244538229.
- Ch10.3: T. Nguyen, "RLHF Methodology," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.3. ISBN: 9798244538229.
- Ch10.5: T. Nguyen, "Human-over-the-Loop," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 10.5. ISBN: 9798244538229.
- Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
- Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note