Observability & Evaluation · Data artifact

Evaluation Rubric

Data artifactObservability & EvaluationObservability & Evaluationarc:EvaluationRubric

A structured scoring specification defining quality criteria (e.g., clarity, completeness, relevance, appropriateness) and their rating scales for human or model evaluators.

Responsibility. Defines consistent scoring criteria for evaluators.

Also known as: Structured rubric, Action Quality Rubric

configuresconfiguresconfiguresLLM Judge: configuresLLM JudgeHuman Evaluator: configuresHuman EvaluatorChain-of-Thought Judge: configuresChain-of-Thought Judge
Direct neighbourhood (hover for relationship types)

Relationships

configures structural

Quantitative guidance

As stated by the sources; verify before use.

Classification

Quality attributes
Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
Risks mitigated
Inconsistent evaluator standards

Sources

  1. Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
  2. Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.
  3. Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
  4. Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
  5. Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note