Observability & Evaluation · Data artifact
Evaluation Rubric
Data artifactObservability & EvaluationObservability & Evaluationarc:EvaluationRubric
A structured scoring specification defining quality criteria (e.g., clarity, completeness, relevance, appropriateness) and their rating scales for human or model evaluators.
Responsibility. Defines consistent scoring criteria for evaluators.
Also known as: Structured rubric, Action Quality Rubric
Relationships
configures structural
Quantitative guidance
As stated by the sources; verify before use.
- Criteria rated clarity 1-5, completeness 1-5, relevance 1-5, appropriateness 1-5 (Ch3.3).
- Tool-selection rubric: intent alignment, parameter appropriateness, efficiency, reasoning soundness, each 1-5 with anchored examples for 5, 3 and 1 (Ch3.8).
- Judge scale: 5 completely correct and comprehensive ... 0 completely wrong or no answer (Ch4.2).
Classification
- Quality attributes
- Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Inconsistent evaluator standards
Sources
- Ch3.3: T. Nguyen, "Web Navigation and Interaction Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.3. ISBN: 9798244538229.
- Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.
- Ch4.2: T. Nguyen, "Deployment and Scaling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 4.2. ISBN: 9798244538229.
- Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
- Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note