Observability & Evaluation · Software component
Response Scorer
Software componentObservability & EvaluationObservability & EvaluationVariation point (abstract)arc:ResponseScorer
An abstract evaluation function that scores one agent response against ground truth, policy, or quality criteria and returns a normalized score or pass/fail indicator.
Responsibility. Converts one agent response into a normalized quality score for aggregation.
Also known as: Custom metric, Scoring function, Evaluator
Variants
| Variant | When to choose |
|---|---|
| Exact Match Scorer | Choose when valid answers have a single canonical form; it fails on correct paraphrases. |
| Fuzzy Match Scorer | Choose when correct answers may be paraphrased (e.g., differently worded delivery dates) but wrong answers must still fail. |
| Keyword Match Scorer | Choose for cheap detection of required/forbidden language (e.g., acknowledgment phrases, unauthorized guarantees). |
| LLM Judge abstract | Choose when the quality dimension is qualitative and resists simple keyword or rule checks. |
| Rule-Based Compliance Scorer | Choose when success depends on structured constraints such as refund limits, dates or eligibility. |
| Semantic Similarity Scorer | Choose when correct answers may be phrased differently from the ground truth and accuracy must be judged by meaning. |
| Task Success Evaluator abstract | — |
| Tool Efficiency Scorer | Choose when the agent exposes a tool-call log and operational efficiency matters. |
| Trajectory Scorer | Choose when deployment requires auditability, interpretability or trust calibration, not just correct outcomes. |
Relationships
is invoked by dependency
reads dependency
sends data to dynamic
Design guidance
- SHOULD be implemented as pure functions that accept the response plus test-case context and return a score normalized to 0-1.
- SHOULD be validated against contrasting examples to confirm it separates good from bad responses.
- SHOULD correlate with business outcomes; metrics that are easy to measure but do not predict value SHOULD NOT drive decisions.
Classification
- Patterns
- Pure scoring functions returning 0-1 scoresPer-case scoring with dataset-level aggregation (mean, percentiles, failure rate)
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Domain-specific failures hidden by generic task success rate
Sources
- Ch3.1A: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1A. ISBN: 9798244538229.
- Ch3.1B: T. Nguyen, "Implement Evaluation Pipelines and Task Benchmarks - Guided Practice," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.1B. ISBN: 9798244538229.