Observability & Evaluation · Software component
LLM Judge
Software componentObservability & EvaluationObservability & EvaluationVariation point (abstract)arc:LLMJudge
A response scorer that prompts a language model with explicit per-level criteria to rate qualitative aspects of an agent response that resist simple rules, returning a normalized score.
Responsibility. Scores qualitative response properties such as empathy by prompting an LLM with a scoring rubric.
Also known as: LLM-as-judge, Model-based grader, LLM-as-a-judge, WebJudge, LLM-as-Judge, Action Quality Judge, VLM-based evaluator, Self-rewarding evaluation mode, Judge LLM, Reasoning judge, LLM evaluator, Hallucination judge
Variant of Response Scorer abstract
When to choose. Choose when the quality dimension is qualitative and resists simple keyword or rule checks.
Variants
| Variant | When to choose |
|---|---|
| AI-Feedback Preference Labeler | Choose when preference labels must scale without human annotation time and be traceable to explicit principles; complement with human labels where contextual nuance matters. |
| Chain-of-Thought Judge | Choose when evaluation accuracy, hallucination detection and user trust matter and secondary verification of the evaluator's logic is needed. |
| Score-Only Judge | Choose only where lower evaluation accuracy and unverifiable judgments are acceptable; the chapter reports it underperforms explained judges. |
Relationships
is configured by structural
invokes dependency
is invoked by dependency
reads dependency
writes dependency
emits telemetry to dynamic
escalates to dynamic
is triggered by dynamic
receives data from dynamic
sends data to dynamic
triggers dynamic
evaluates assurance
is evaluated by assurance
alternative to variability
Design guidance
- SHOULD provide detailed criteria for each score level when prompting the judge model.
- SHOULD have edge-case disagreements validated by human evaluators.
- SHOULD be adversarially tested with known-incorrect outputs, since agents may learn to optimize for the judge rather than the task.
- SHOULD use an evaluator model stronger than the agent under evaluation.
- SHOULD average scores from multiple evaluators and calibrate rubrics against human validation to control variance.
- SHOULD supplement rather than replace human feedback when used for self-rewarding, because models may score their own outputs highly regardless of quality (reward hacking).
- SHOULD compare candidate judge models to find the best fit for the domain.
- MUST use explicit rubrics that turn vague quality concepts into concrete, answerable criteria.
- SHOULD require the judge to explain its ratings so humans can verify judgments and agents receive actionable feedback.
- SHOULD be periodically audited by human experts to detect drift or systematic bias, especially in domains requiring deep expertise.
- SHOULD be reserved for high-risk applications or outputs flagged by earlier stages due to latency and cost.
- MUST use a specific rubric with anchored score levels rather than asking 'is this hallucinated?'.
- SHOULD run at temperature 0 with structured (JSON) output of score, reasoning, specific hallucinations, severity and recommendation.
- SHOULD NOT be relied on alone, since judges can themselves hallucinate in evaluations.
- SHOULD score traces against an explicit rubric (e.g., accuracy, completeness, clarity on 1-5) for scalable, consistent automated evaluation.
- SHOULD be validated by human review because it depends on judge-model quality, may miss subtle errors and handles subjective quality poorly.
Quantitative guidance
As stated by the sources; verify before use.
- Example judge scale: empathy rated 1-5 (Ch3.1B).
- LLM judges reach ~85% agreement with human judgment, implying ~15% systematic disagreement (Ch3.3).
- WebJudge: ~85% human agreement with a 3.8% average success-rate gap versus human evaluation (Ch3.3).
- RAGAS evaluator judge LLM configured with an 8-token maximum output; recommended judges ranked Mixtral 8x22B, Mixtral 8x7B, Llama 3.1 70B, Llama 3.3 70B (Ref3.01).
Classification
- Patterns
- LLM-as-judgeLLM-as-a-judgeProgress assessment from screenshots and actionsMulti-evaluator score averagingHuman-calibrated rubricsSelf-rewarding language modelsScalable first-pass evaluation with human review of flagged casesDefense-in-depth stage 4 (sampling)Rubric-based judgementConstitutional AI as principle-criteria LLM-as-Judge
- Technologies
- GPT-4WebJudgeClaudeNVIDIA NeMo EvaluatorMixtral 8x22B InstructMixtral 8x7B InstructLlama 3.1 70B InstructLlama 3.3 70B Instruct
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)Explainability (NIST AI RMF: explainable and interpretable)
- Risks mitigated
- Qualitative quality gaps invisible to accuracy metricsRule-based checks penalizing valid alternative pathsFalse negatives of programmatic validation when multiple valid tools or equivalent values existReliance on subjective human judgment aloneAnnotation bottlenecksEvaluation that resists automated metrics lacking ground truthSubtle semantic hallucinations missed by deterministic and statistical checksVerification failures masquerading as successes
Sources
- Ch3.1B: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks - Guided Practice," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1B. ISBN: 9798244538229.
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
- Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.
- Ch3.9: T. Nguyen, "Reasoning Quality," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.9. ISBN: 9798244538229.
- Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
- Ch4.1: T. Nguyen, "Introduction to AI Agent Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.1. ISBN: 9798244538229.
- Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
- Ch5.8: T. Nguyen, "Semantic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.8. ISBN: 9798244538229.
- Ch8.2A: T. Nguyen, "Error Rates and Reliability," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.2A. ISBN: 9798244538229.
- Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
- Ref3.01: NVIDIA, "Agent Evaluation in NVIDIA NeMo Agent Toolkit," NVIDIA NeMo Agent Toolkit Documentation, v1.8. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/agent-toolkit/latest/improve-workflows/evaluate.html
- Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
- Ref8.02: "Machine Learning Monitoring in Production," unpublished reference note (02-ML-Monitoring-Production.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note