Observability & Evaluation · Software component
Tool Call Accuracy Evaluator
Software componentObservability & EvaluationObservability & Evaluationarc:ToolCallAccuracyEvaluator
A programmatic evaluator that validates recorded tool calls against schemas and ground-truth references, scoring tool selection, per-parameter correctness and execution success.
Responsibility. Deterministically scores the correctness of individual tool calls.
Also known as: Programmatic Tool Call Validator, Function Calling Evaluator, Tool calling correctness assessment
Relationships
is invoked by dependency
reads dependency
emits telemetry to dynamic
evaluates assurance
- Agent Controller abstract Ch3.8
Design guidance
- SHOULD report tool selection, parameter (by type) and execution success separately rather than a single aggregate.
- SHOULD use strict matching for security-critical parameters and lenient semantic matching for user-facing values.
- SHOULD distinguish tool-side (system) failures from agent choice or parameter failures.
- SHOULD measure tool appropriateness, parameter correctness against context and the rate of invalid tool calls (Ref8.02, Ref8.03).
Quantitative guidance
As stated by the sources; verify before use.
- Tool selection > 90% indicates good tool semantics understanding; < 75% needs better descriptions/examples; simple tasks 95-98%, ambiguous tasks 75-85% (Ch3.8).
- Example hierarchy: overall 84.2%, tool selection 95.8%, parameters 78.4% (temporal 58.9%), execution 89.3%, error recovery 51.6% (Ch3.8).
- Example diagnosis: function name accuracy 92%, parameter hallucination 18%, success rate 76% (Ch3.8).
Classification
- Patterns
- Format, parameter and execution validationStrict vs lenient parameter matchingHierarchical metric reporting by action and parameter typeBFCL metrics (function name accuracy, parameter hallucination rate, parameter missing rate, progress rate, success rate)Session-level plus node-level metrics
- Technologies
- Berkeley Function Calling Leaderboard
- Quality attributes
- Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Maintainability (ISO/IEC 25010)
- Risks mitigated
- Averaging away diagnostic informationConflating tool selection with tool calling accuracyMisattributing system failures to agent competence
Sources
- Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.
- Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
- Ref8.02: "Machine Learning Monitoring in Production," unpublished reference note (02-ML-Monitoring-Production.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
- Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note