Safety & Security · Software component
Tool Hallucination Detector
Software componentSafety & SecuritySafety, Security & GovernanceVariation point (abstract)arc:ToolHallucinationDetector
An abstract detector that estimates, from model confidence signals, whether a proposed tool call or parameter is likely hallucinated, flagging low-confidence calls for verification.
Responsibility. Scores the likelihood that a proposed tool call is fabricated so suspicious calls can be verified before execution.
Also known as: Confidence-based tool call verification
Variants
| Variant | When to choose |
|---|---|
| Entropy-Based Hallucination Detector | Choose when the generating model's token probability distributions are observable during parameter generation, so uncertainty can be measured without a separate model. |
| Verifier-Model Hallucination Detector | Choose when labeled examples of correct and hallucinated tool calls are available to train a well-calibrated verifier, and a second model check per call is acceptable. |
Relationships
Design guidance
- SHOULD require explicit confirmation before proceeding when confidence falls below a threshold.
- SHOULD NOT substitute for improving tool documentation, which often reduces hallucination rates more effectively than detection mechanisms.
Classification
- Patterns
- Uncertainty-based hallucination detection
- Quality attributes
- Safety (ISO/IEC 25010 | NIST AI RMF: safe)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- Tool hallucination (invented tools or fabricated parameters)
Sources
- Ch3.7: T. Nguyen, "Tool Usage Auditing," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.7. ISBN: 9798244538229.