Observability & Evaluation · Software component

Tool Call Accuracy Evaluator

Software componentObservability & EvaluationObservability & Evaluationarc:ToolCallAccuracyEvaluator

A programmatic evaluator that validates recorded tool calls against schemas and ground-truth references, scoring tool selection, per-parameter correctness and execution success.

Responsibility. Deterministically scores the correctness of individual tool calls.

Also known as: Programmatic Tool Call Validator, Function Calling Evaluator, Tool calling correctness assessment

evaluatesis invoked byemits telemetry toreadsreadsreadsAgent Controller: evaluatesAgent ControllerEvaluation Harness: is invoked byEvaluation HarnessMetrics Collector: emits telemetry toMetrics CollectorTrace Store: readsTrace StoreTool Schema: readsTool SchemaReference Trajectory: readsReference Trajectory
Direct neighbourhood (hover for relationship types)

Relationships

is invoked by dependency

reads dependency

emits telemetry to dynamic

evaluates assurance

Design guidance

Quantitative guidance

As stated by the sources; verify before use.

Classification

Patterns
Format, parameter and execution validationStrict vs lenient parameter matchingHierarchical metric reporting by action and parameter typeBFCL metrics (function name accuracy, parameter hallucination rate, parameter missing rate, progress rate, success rate)Session-level plus node-level metrics
Technologies
Berkeley Function Calling Leaderboard
Quality attributes
Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)Maintainability (ISO/IEC 25010)
Risks mitigated
Averaging away diagnostic informationConflating tool selection with tool calling accuracyMisattributing system failures to agent competence

Sources

  1. Ch3.8: T. Nguyen, "Action Accuracy Metrics," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.8. ISBN: 9798244538229.
  2. Ref8.01: LangChain, "LangSmith observability: AI agent observability platform," LangChain. Accessed: Sep. 27, 2026. [Online]. Available: https://www.langchain.com/langsmith/observability
  3. Ref8.02: "Machine Learning Monitoring in Production," unpublished reference note (02-ML-Monitoring-Production.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note
  4. Ref8.03: "Agent Evaluation Frameworks and Metrics," unpublished reference note (03-Agent-Evaluation-Frameworks.md), Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam supplementary materials, 2026. unpublished note