Observability & Evaluation · Software component
Semantic Similarity Scorer
Software componentObservability & EvaluationObservability & Evaluationarc:SemanticSimilarityScorer
A response scorer that measures accuracy as the semantic similarity between an agent response and an expert-validated ground-truth answer.
Responsibility. Scores response accuracy by meaning-level similarity to ground truth.
Also known as: Semantic accuracy metric
Variant of Response Scorer abstract
When to choose. Choose when correct answers may be phrased differently from the ground truth and accuracy must be judged by meaning.
Relationships
is invoked by dependency
Design guidance
- SHOULD score against ground-truth answers curated and validated by domain experts.
Quantitative guidance
As stated by the sources; verify before use.
- Example per-test semantic similarity 0.87-0.92 against an 85% accuracy threshold (Ch7.3).
Classification
- Patterns
- Ground-truth benchmarking
- Technologies
- NVIDIA NeMo Agent Toolkit Evaluator
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Reliability (ISO/IEC 25010 | NIST AI RMF: valid and reliable)
- Risks mitigated
- Silent accuracy regressions after prompt or model changes
Sources
- Ch7.3: T. Nguyen, "NeMo Agent Toolkit Profiling," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 7.3. ISBN: 9798244538229.