Safety & Security · Software component
Verifier-Model Hallucination Detector
Software componentSafety & SecuritySafety, Security & Governancearc:VerifierModelHallucinationDetector
A tool hallucination detector that submits each proposed tool call to a separately trained, calibrated verifier model and allows execution only when the verifier confirms the call is valid.
Responsibility. Validates primary-model tool calls with an independent verifier model's calibrated confidence.
Also known as: Confidence calibration, Two-model architecture, Secondary verifier
Variant of Tool Hallucination Detector abstract
When to choose. Choose when labeled examples of correct and hallucinated tool calls are available to train a well-calibrated verifier, and a second model check per call is acceptable.
Relationships
invokes dependency
alternative to variability
Design guidance
- SHOULD proceed only when the verifier confirms the primary model's tool call makes sense.
Classification
- Patterns
- Generator-verifier (two-model) architectureConfidence calibration
- Quality attributes
- Safety (ISO/IEC 25010 | NIST AI RMF: safe)Functional suitability: correctness and validity (ISO/IEC 25010 | NIST AI RMF: valid)
- Risks mitigated
- Hallucinated tool invocations passing syntactic checks
Sources
- Ch3.7: T. Nguyen, "Tool Usage Auditing," in Mastering Agentic AI Systems: Guide for the NVIDIA NCP-AAI Exam, 1st ed. 2026, ch. 3.7. ISBN: 9798244538229.