Observability & Evaluation · Software component

LLM Judge

Software componentObservability & EvaluationObservability & EvaluationVariation point (abstract)arc:LLMJudge

A response scorer that prompts a language model with explicit per-level criteria to rate qualitative aspects of an agent response that resist simple rules, returning a normalized score.

Responsibility. Scores qualitative response properties such as empathy by prompting an LLM with a scoring rubric.

Also known as: LLM-as-judge, Model-based grader, LLM-as-a-judge, WebJudge, LLM-as-Judge, Action Quality Judge, VLM-based evaluator, Self-rewarding evaluation mode, Judge LLM, Reasoning judge, LLM evaluator, Hallucination judge

Variant of Response Scorer abstract

When to choose. Choose when the quality dimension is qualitative and resists simple keyword or rule checks.

evaluates; triggersescalates to; is evaluated byevaluates; is configured byreceives data from; is triggered byevaluatesinvokesis invoked bywritesemits telemetry toevaluatesreadsis invoked byevaluatesis triggered byis invoked bywritesspecializessends data toAnswer Synthesizer: evaluates; triggersAnswer SynthesizerHuman Evaluator: escalates to; is evaluated byHuman EvaluatorPrompt Exemplar Set: evaluates; is configured byPrompt Exemplar SetEvaluation Trace Sampler: receives data from; is triggered byEvaluation Trace SamplerAgent Controller: evaluatesAgent ControllerLLM Inference Service: invokesLLM Inference ServiceEvaluation Harness: is invoked byEvaluation HarnessAudit Log Store: writesAudit Log StoreMetrics Collector: emits telemetry toMetrics CollectorReasoning Engine: evaluatesReasoning EngineTrace Store: readsTrace StoreOnline Evaluator: is invoked byOnline EvaluatorWeb Navigation Agent: evaluatesWeb Navigation AgentCitation Verifier: is triggered byCitation VerifierExperiment Guardrail Monitor: is invoked byExperiment Guardrail Mon…Preference Dataset: writesPreference DatasetResponse Scorer: specializesResponse ScorerEvaluation Failure Analyzer: sends data toEvaluation Failure Analy…+18 more (see relationships)
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
AI-Feedback Preference LabelerChoose when preference labels must scale without human annotation time and be traceable to explicit principles; complement with human labels where contextual nuance matters.
Chain-of-Thought JudgeChoose when evaluation accuracy, hallucination detection and user trust matter and secondary verification of the evaluator's logic is needed.
Score-Only JudgeChoose only where lower evaluation accuracy and unverifiable judgments are acceptable; the chapter reports it underperforms explained judges.

Relationships

is configured by structural

invokes dependency

is invoked by dependency

reads dependency

writes dependency

emits telemetry to dynamic

escalates to dynamic

is triggered by dynamic

receives data from dynamic

sends data to dynamic

triggers dynamic

evaluates assurance

is evaluated by assurance

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
LLM-as-judgeLLM-as-a-judgeProgress assessment from screenshots and actionsMulti-evaluator score averagingHuman-calibrated rubricsSelf-rewarding language modelsScalable first-pass evaluation with human review of flagged casesDefense-in-depth stage 4 (sampling)Rubric-based judgementConstitutional AI as principle-criteria LLM-as-Judge
Technologies
GPT-4WebJudgeClaudeNVIDIA NeMo EvaluatorMixtral 8x22B InstructMixtral 8x7B InstructLlama 3.1 70B InstructLlama 3.3 70B Instruct
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Performance efficiency (ISO/IEC 25010)Explainability (NIST AI RMF: explainable and interpretable)
Risks mitigated
Qualitative quality gaps invisible to accuracy metricsRule-based checks penalizing valid alternative pathsFalse negatives of programmatic validation when multiple valid tools or equivalent values existReliance on subjective human judgment aloneAnnotation bottlenecksEvaluation that resists automated metrics lacking ground truthSubtle semantic hallucinations missed by deterministic and statistical checksVerification failures masquerading as successes

Sources

  1. Ch3.1B: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks - Guided Practice," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1B. ISBN: 9798244538229.
  2. Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
  3. Ch3.5: T. Nguyen, "Prompt Optimization, Few-Shot Learning, Fine-Tuning," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.5. ISBN: 9798244538229.
  4. Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.
  5. Ch3.9: T. Nguyen, "Reasoning Quality," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.9. ISBN: 9798244538229.
  6. Ch3.10: T. Nguyen, "Efficiency Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.10. ISBN: 9798244538229.
  7. Ch4.1: T. Nguyen, "Introduction to AI Agent Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.1. ISBN: 9798244538229.
  8. Ch4.4: T. Nguyen, "Performance Profiling and Optimization," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.4. ISBN: 9798244538229.
  9. Ch5.8: T. Nguyen, "Semantic Memory," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 5.8. ISBN: 9798244538229.
  10. Ch8.2A: T. Nguyen, "Error Rates and Reliability," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 8.2A. ISBN: 9798244538229.
  11. Ch9.5: T. Nguyen, "Constitutional AI," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 9.5. ISBN: 9798244538229.
  12. Ref3.01: NVIDIA, "Agent Evaluation in NVIDIA NeMo Agent Toolkit," NVIDIA NeMo Agent Toolkit Documentation, v1.8. Accessed: Sep. 27, 2026. [Online]. Available: https://docs.nvidia.com/nemo/agent-toolkit/latest/improve-workflows/evaluate.html
  13. Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
  14. Ref8.02: "Machine Learning Monitoring in Production," unpublished reference note (02-ML-Monitoring-Production.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  15. Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note