Observability & Evaluation · Software component

Chain-of-Thought Judge

Software componentObservability & EvaluationObservability & Evaluationarc:ChainOfThoughtJudge

An LLM judge that produces an explicit reasoning trace justifying its evaluation before giving its verdict.

Responsibility. Grades outputs with a visible, verifiable evaluation rationale.

Also known as: CoT Reasoning Agent evaluator, Explained-rating judge, Correctness evaluator, Completeness evaluator, Safety evaluator

Variant of LLM Judge abstract

When to choose. Choose when evaluation accuracy, hallucination detection and user trust matter and secondary verification of the evaluator's logic is needed.

is invoked byspecializesis invoked byis configured byalternative toEvaluation Harness: is invoked byEvaluation HarnessLLM Judge: specializesLLM JudgeOnline Evaluator: is invoked byOnline EvaluatorEvaluation Rubric: is configured byEvaluation RubricScore-Only Judge: alternative toScore-Only Judge
Direct neighbourhood (hover for relationship types)

Relationships

is configured by structural

is invoked by dependency

alternative to variability

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Risks mitigated
Rating verbose but incorrect answers highlyPenalising concise correct answers

Sources

  1. Ch3.9: T. Nguyen, "Reasoning Quality," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.9. ISBN: 9798244538229.
  2. Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.