Observability & Evaluation · Software component

Response Scorer

Software componentObservability & EvaluationObservability & EvaluationVariation point (abstract)arc:ResponseScorer

An abstract evaluation function that scores one agent response against ground truth, policy, or quality criteria and returns a normalized score or pass/fail indicator.

Responsibility. Converts one agent response into a normalized quality score for aggregation.

Also known as: Custom metric, Scoring function, Evaluator

is invoked byis specialized byreadsis specialized byis specialized byis specialized bysends data tois specialized byis specialized byis specialized byis specialized byis specialized byEvaluation Harness: is invoked byEvaluation HarnessLLM Judge: is specialized byLLM JudgeEvaluation Dataset: readsEvaluation DatasetTask Success Evaluator: is specialized byTask Success EvaluatorKeyword Match Scorer: is specialized byKeyword Match ScorerTrajectory Scorer: is specialized byTrajectory ScorerEvaluation Score Aggregator: sends data toEvaluation Score Aggrega…Tool Efficiency Scorer: is specialized byTool Efficiency ScorerExact Match Scorer: is specialized byExact Match ScorerFuzzy Match Scorer: is specialized byFuzzy Match ScorerRule-Based Compliance Scorer: is specialized byRule-Based Compliance Sc…Semantic Similarity Scorer: is specialized bySemantic Similarity Scorer
Direct neighbourhood (hover for relationship types)

Variants

VariantWhen to choose
Exact Match ScorerChoose when valid answers have a single canonical form; it fails on correct paraphrases.
Fuzzy Match ScorerChoose when correct answers may be paraphrased (e.g., differently worded delivery dates) but wrong answers must still fail.
Keyword Match ScorerChoose for cheap detection of required/forbidden language (e.g., acknowledgment phrases, unauthorized guarantees).
LLM Judge abstractChoose when the quality dimension is qualitative and resists simple keyword or rule checks.
Rule-Based Compliance ScorerChoose when success depends on structured constraints such as refund limits, dates or eligibility.
Semantic Similarity ScorerChoose when correct answers may be phrased differently from the ground truth and accuracy must be judged by meaning.
Task Success Evaluator abstract—
Tool Efficiency ScorerChoose when the agent exposes a tool-call log and operational efficiency matters.
Trajectory ScorerChoose when deployment requires auditability, interpretability or trust calibration, not just correct outcomes.

Relationships

is invoked by dependency

reads dependency

sends data to dynamic

Design guidance

Classification

Patterns
Pure scoring functions returning 0-1 scoresPer-case scoring with dataset-level aggregation (mean, percentiles, failure rate)
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Maintainability (ISO/IEC 25010)
Risks mitigated
Domain-specific failures hidden by generic task success rate

Sources

  1. Ch3.1A: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1A. ISBN: 9798244538229.
  2. Ch3.1B: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks - Guided Practice," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1B. ISBN: 9798244538229.